The story everyone told was “the AI mispriced houses.” That is the shallow reading. The model was never the failure. The disabled override was, and it is the exact decision on every agent team's whiteboard right now.
Early in 2021, Zillow made a decision that, in hindsight, reads less like a pricing strategy and more like a controlled demolition of its own immune system. The company ran a home-flipping business, Zillow Offers, that bought houses directly from sellers, held them briefly, and resold them. The engine behind the offers was the Zestimate, Zillow's famous algorithmic home-value estimate, and the company had a team of human pricing experts whose job was to sanity-check what the model spat out. Under an initiative reported internally as “Project Ketchup,” Zillow did two things at once. It began using the Zestimate directly as its cash offer on qualifying homes. And, according to business-press reporting on the program, it prevented those pricing experts from modifying the algorithm's valuations and asked them to stop questioning them.
The experts did not leave. Their desks did not move. The company simply told them that the number the model produced was the number, full stop, and their job was no longer to argue with it. Acquisition volumes did exactly what you would expect once the brakes were disconnected: they reportedly more than doubled in a single quarter. Zillow was buying houses faster than it ever had, at prices no human was allowed to override downward.
By November 2021, it was over. Zillow announced it was winding down Zillow Offers entirely and cutting about a quarter of its workforce, roughly 2,000 people. For the full year ended December 31, 2021, its 10-K recorded, in the filing's own language, “a write-down to inventory totaling $407.9 million” as a result of “unintentionally purchasing homes at higher prices than the Company's current estimates of the future selling prices.” Four hundred and seven point nine million dollars of houses bought for more than they were worth, by a system whose one human correction channel had been switched off on purpose, at the worst possible moment to switch it off.
The story everyone told afterward was “the AI mispriced houses.” That story is not wrong, exactly, but it is the shallow reading, and the shallow reading buries the lesson that actually transfers to anyone shipping automated decisions in 2026. Because the model was never the failure. The disabled override was.
Here is the thing the “the AI failed” framing misses: a pricing model being wrong sometimes is not a defect. It is the baseline condition of pricing models, and every serious operation that uses one prices that in. The Zestimate had a known error distribution. On the vast majority of homes it was close enough, and on some homes, the unusual ones, the fast-moving markets, the properties with quirks the training data underrepresented, it was off, sometimes badly. This was not a secret. It was the whole reason a team of human experts existed in the first place. Their value was never in the 95% of cases where the model was right, where they were pure overhead. Their value was entirely in the 5% where it was wrong, and specifically in catching the wrong ones before Zillow wired the money.
What Project Ketchup did was delete the mechanism that caught the 5%, in exchange for the speed of trusting the 95%. And that trade looks brilliant right up until the distribution shifts, at which point the 5% stops being a scattered, tolerable error rate and becomes a correlated, systemic one. In 2021, the U.S. housing market did something very few forecasters called, moving in ways that broke the recent-history assumptions baked into automated valuation. The model did not get suddenly stupid. The ground moved under it, its errors stopped canceling out and started stacking in one direction, and the one system that could have noticed, “these offers have been running hot for weeks, something's off”, had been told to stop noticing. The experts were still in the building. They had been converted from a correction channel into spectators.
Read the CEO's own explanation with this in mind, because it is more precise than the coverage gave it credit for. Rich Barton said the company had “determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility.” Notice what that sentence is actually confessing. It is not a confession about accuracy, about the model's central estimate being biased. It is a confession about variance. The problem was not that the Zestimate was consistently wrong; it was that its errors had a spread the business could not absorb, and Zillow had spent the year removing every shock absorber it had. Barton is describing a company that scaled up its exposure to a fat tail while dismantling the thing that clipped the tail. The word doing the work in his statement is “volatility,” and volatility is precisely what human review exists to dampen.
There is a temptation, writing about a forecasting failure, to be sloppy with the very numbers whose sloppiness is the subject, and this story is a minefield for it, because at least four different dollar figures circulate as “the Zillow number” and they refer to four different things. The $407.9 million above is the full-year inventory write-down from the 10-K. A separate, widely-cited $421 million is the Q3 2021 loss of the iBuying segment, a different quantity measuring a different thing. There was a $304 million write-down reported for Q3 specifically, and a forward-looking estimate of another $240 to $265 million in losses expected in Q4. A larger round number, often quoted as the “total,” floats around secondary coverage without a clean line-item behind it, and a careful writer simply does not print it, because a postmortem that inherits an unsourced figure is committing, in miniature, the exact error it is diagnosing.
And there is a small, sharp irony worth pausing on, visible only if you keep the numbers straight. Add the Q3 write-down to the Q4 forecast, $304 million plus the $240 to $265 million expected, and you get a projected inventory loss somewhere in the range of $544 to $569 million. The actual full-year write-down came in at $407.9 million. The company's own forecast of its losses overshot the reality by well over a hundred million dollars. A firm brought down by the unpredictability of its forecasts also could not accurately forecast the size of its own failure, in the optimistic direction. This is not a gotcha; it is the same lesson wearing a different suit. Forecasting is hard, the tails are wide, and a number stated with confidence is not the same as a number that came true, whether the number is a home price or a projected loss. If you are going to write about a company that trusted its predictions too much, the least you can do is hold your own predictions loosely.
One more piece of honesty the shallow version skips. This is not evidence that algorithmic home pricing is doomed, and it is not evidence that Zillow's engineers were fools. iBuying as a model survived Zillow's exit; competitors kept operating. And the crucial distinction is between two kinds of wrong. The forecast was wrong ex post, after the fact, once the market did its unlikely thing, and blaming anyone for not predicting an unpredictable market is cheap hindsight. The governance decision, disabling the override to gain speed, was wrong ex ante, wrong at the moment it was made, regardless of how the market turned out, because it traded away the ability to respond to being wrong. You cannot fault Zillow for failing to see the future. You can fault it for deliberately removing its own capacity to react when the future arrived, which is a decision that looked correct only because it had not yet been tested by a bad draw.
Here is why this is a 2026 story and not a 2021 one. Strip away the houses and the Zestimate, and Zillow's decision is the exact decision on the whiteboard at every company shipping AI agents this year. You have a model. It is right most of the time. You have some human-in-the-loop review, an approval step, an expert who can veto or amend what the model proposes. And that review is slow, and it is expensive, and it visibly does not scale, and someone in the room can produce a chart showing that the humans agree with the model the overwhelming majority of the time, so what, exactly, are we paying them for? The pressure to remove the override runs in one direction, always, because the cost of the override is a line item you can see and the cost of removing it is a tail you cannot see until it arrives.
Zillow is the priced version of that argument, and the price was $407.9 million and a business. The reason the review looked like pure cost is the same reason it was not: when your model is right 95% of the time, the human check appears to add nothing 95% of the time, and earns its entire annual keep in a handful of cases you cannot identify in advance. You are not paying for the average case. You are paying for the correlated bad quarter, the distribution shift, the moment the model's errors line up and start pointing the same way. Removing the override is deleting insurance because you have not had a claim, at exactly the point in the cycle where the claim is coming.
And notice how the correction would actually have worked, because this is the part that makes the override cheap and the loss expensive. No one needed to catch each individual mispriced house; that is genuinely infeasible at Zillow's volume, and it is the fair case for automating. What a human channel catches is the aggregate signal, the thing no single transaction reveals but a person watching the flow can see: offers running consistently above eventual sale prices for weeks, acquisition volume spiking while margins quietly invert, the smell of a book that is filling with homes bought too high. That is a slow, boring, one-analyst-with-a-dashboard job, and it is exactly the job Project Ketchup defined out of existence when it told the experts the algorithm's number was final. The correlated error announces itself in the aggregate long before it lands in the write-down. Zillow removed the only role positioned to hear it.
So the practical rule to carry out of Zillow is not “never automate” or “always keep a human in the loop,” both too blunt to be useful. It is narrower and more actionable than that. Before you disable a correction channel, ask what happens to your exposure when the model is wrong not randomly but systematically, all in the same direction at once, because that is the failure the override exists to catch, and it is invisible in every metric computed during good times. And treat the argument “the model is usually right, so the review is overhead” as a red flag rather than a business case, because that sentence is not describing overhead. It is describing insurance, and the word “usually” is doing all the work: it is a precise measurement of how often you will wish you had kept the thing you are about to delete. Zillow had the experts. It had the model. It had the money. What it removed, to go faster, was the one cheap mechanism that stood between a routine model error and a nine-figure write-down, and it removed it in the calm before the exact storm that mechanism was for.
A companion piece on this blog once looked at a company gutted from the outside, its business model destroyed by someone else's product. Zillow is the mirror image, a company gutted from the inside, by its own model, with the humans who could have stopped it standing right there, told to watch. The outside kind you cannot always prevent. The inside kind you do to yourself, and it always looks, on the day you do it, like progress.
“The model is usually right, so the review is overhead” is not a business case. It is a description of insurance.
Zillow's loss was a governance failure: it disabled the one channel that could have caught its model's errors lining up in one direction. Every team shipping AI agents faces the same whiteboard decision, and the correlated failure the override exists to catch is invisible in every good-times metric. The Agent Trust Stack is the machinery for keeping that correction channel wired: provenance to see what each agent actually did, verification gates that stay in the loop by design, and reputation so drift shows up in the aggregate before it lands in a write-down. Keep the shock absorber; watch the flow, not just the average case.
Read the Theory of Agent Trust
Install the whole stack: pip install agent-trust-stack · npm install agent-trust-stack