← Back to blog

"The Year of the Agent" Was the Year of the Demo

The letter of the prediction squeaked by. The spirit missed. And the year earned a more precise name than anyone gave it at the time.

Published July 2026 · 10 min read · agent reliability / prediction accountability / enterprise AI / 2025 retrospective


In February 2024, the Swedish fintech Klarna announced that its new AI assistant, built with OpenAI, had handled 2.3 million customer service conversations in its first month, cut average resolution time from eleven minutes to under two, and was doing the work of 700 human agents. The company projected $40 million in profit improvement from it that year. It was the cleanest possible specimen of the thing the industry was about to spend eighteen months promising: an agent had, apparently, joined the workforce, and you could count the headcount it replaced.

By May 2025, Klarna was hiring customer service humans again. CEO Sebastian Siemiatkowski said it himself, in an unusually honest series of interviews: cost had been allowed to dominate the evaluation, and what the company ended up with was, in his words, lower quality. The assistant still handles plenty of volume. But the arc, from "700 agents' worth of work" to "we went too far, quality suffered, humans are coming back," is the story of the agent era so far, compressed into one company and fifteen months. (We took apart what those Klarna numbers actually measured separately; here it is the opening specimen, not the subject.)

Right in the middle of that arc, on January 6, 2025, Sam Altman published a short blog post called "Reflections," and one sentence from it became the banner over the entire year: "We believe that, in 2025, we may see the first AI agents 'join the workforce' and materially change the output of companies."

The year is over. We can score the sentence now. And scoring it honestly turns out to be much more interesting than dunking on it.

Score the sentence he wrote, not the one you remember

Read Altman's sentence again, slowly, because it is more hedged than its reputation. "May see," not will. The "first" AI agents, not a wave. And "join the workforce" is in scare quotes in the original, doing the work of a metaphor, not a census claim. The sentence the internet remembers, agents will replace workers in 2025, is not the sentence he wrote. If this essay is about prediction accountability, the first rule is to hold the man to his actual words, because flattening a hedged claim into a bold one and then debunking the bold one is its own kind of hallucination.

The context, though, tells you the spirit. The sentence immediately before it reads: "We are now confident we know how to build AGI as we have traditionally understood it." OpenAI's chief product officer, Kevin Weil, glossed the same prediction as the year "that we go from ChatGPT being this super smart thing... to ChatGPT doing things in the real world for you." And nobody at any vendor spent 2025 talking down expectations. The letter of the claim was modest. The spirit, the one the industry funded, was transformation.

So score both. The letter says: some first agents might materially change the output of some companies. The spirit says: this is the year the needle moves. One of those survives contact with the data. The other does not.

What the receipts say

The most consequential accounting of the year came from MIT's Project NANDA, whose "State of AI in Business 2025" report reviewed more than 300 publicly disclosed AI initiatives, interviewed people at 52 organizations, and surveyed 153 senior leaders. Its headline finding, stated in its own executive summary: "Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return." Just 5 percent of integrated pilots, the report found, were extracting real value. The report's authors named the gulf the GenAI Divide. A fair caveat belongs here: the study's methodology drew real criticism, and one report is one report. But no serious competing accounting of 2025 found the opposite, and the direction matches everything else on this list.

What I find most damning is not the 95 percent. It is a single anonymous CIO quoted inside the report: "We've seen dozens of demos this year. Maybe one or two are genuinely useful. The rest are wrappers or science projects." That is the year of the agent as experienced from the buying side. Demos arrived in dozens. Deployments arrived in ones and twos.

The rest of the receipts stack the same way. Gartner, in a June 25, 2025 press release, predicted that over 40 percent of agentic AI projects would be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls, and estimated that of the thousands of vendors claiming to sell agentic AI, only about 130 actually were. Gartner even coined a term for the gap: agent washing, the rebranding of chatbots and RPA as agents. Meanwhile, the most sobering benchmark of the era came from Carnegie Mellon, whose TheAgentCompany placed agents inside a simulated software company and assigned them 175 real professional tasks: writing code, browsing internal tools, coordinating with simulated colleagues. The best agent of the initial evaluation, running on Claude 3.5 Sonnet, autonomously completed 24 percent of tasks. Later frontier models pushed that to roughly 30 percent. In the most charitable reading, the state of the art failed seven out of ten office tasks in the very year it was supposed to join the office.

And yet the honest scorecard has a second column. Coding agents genuinely worked, and reshaped how a lot of software got written. Deep-research agents genuinely worked. And MIT's own 5 percent were not lottery winners: the report found that external partnerships with learning-capable, customized tools reached deployment about 67 percent of the time, versus about 33 percent for internally built generic ones, a self-reported figure the report itself flags, but a consistent one. The report even found a thriving "shadow AI economy": employees quietly using personal ChatGPT and Claude accounts to automate chunks of their own jobs while their employers' official initiatives stalled in pilot. Read that carefully, because it is the letter of Altman's sentence coming true in the least expected place. Some agents did join the workforce in 2025. They were smuggled in by the workers.

So the verdict: the letter squeaks by, on the 5 percent, on the coding agents, on the shadow economy. The spirit missed. And the year earned a more precise name than anyone gave it at the time. Demos were everywhere; dozens per CIO. The workforce, officially, barely changed. The year of the agent was the year of the demo.

Why the demo works and the deployment doesn't

Here is the part I care about most, because it is not a story about hype or dishonesty. There is a mechanical reason the same agent dazzles in a demo and dies in production, and you can compute it on a napkin.

A demo is short. A production workflow is long. Suppose your agent executes any single step correctly 85 percent of the time, a number that feels impressively high when you watch it work. A three-step demo succeeds 0.85 cubed, about 61 percent of the time: flaky, but you run it twice for the customer and it shines. A fifteen-step unattended workflow succeeds 0.85 to the fifteenth power: about 9 percent of the time. Same agent. Same per-step competence. One is a demo you would fund. The other is a production incident with a login page. I have written before about this compounding wall, the 85 percent problem, and 2025 was its coming-out year: the demo is not lying to you, it is simply short enough to survive its own error rate.

Look at MIT's named failure mechanism through that lens. The report's diagnosis for why pilots stall is what it calls the learning gap: "tools that don't learn, integrate poorly, or match workflows," systems that do not retain feedback or improve over time, with users abandoning them for mission-critical work "due to lack of memory." Every one of those is a per-step error source that compounding turns lethal. A tool that misreads your context occasionally is charming across three steps and catastrophic across thirty. "Join the workforce" is not a bigger demo. It is a categorically different reliability regime, and 2025 cleared the demo bar and not the other one.

The person who put this most bluntly is not a critic on the outside but an architect on the inside. Andrej Karpathy, an OpenAI co-founder who uses these tools every day, said in late 2025 that when he hears "2025 is the year of agents" he gets concerned, because in his framing this is the decade of agents, not the year. His reason is the compounding wall named from within the lab: agents, he said, do not have continual learning, you cannot just tell them something and have them remember it. That is MIT's lack-of-memory failure mode and the per-step error source above, described by one of the people building the things. He is not saying agents do not matter; he uses them daily. He is saying the prediction had the right verb and the wrong exponent, off by roughly a factor of ten on the timeline. Cal Newport reached the same verdict from the outside, scoring Altman's sentence as a claim that "ended up not happening" while noting that we do not really know how to build the digital employees we were promised. The compounding arithmetic is the missing why underneath both of them: it is not a mystery that the digital employees did not arrive, it is multiplication.

Which is why "give it another year" is the wrong lesson, and so is "it was all vapor." The right lesson is a century old.

The dynamo in the demo room

In 1987, the economist Robert Solow quipped that you could see the computer age everywhere but in the productivity statistics. Three years later, the economic historian Paul David published a famous explanation, "The Dynamo and the Computer," about the last time this happened. Electric motors were demonstrated, commercialized, and celebrated in the 1880s. American factory productivity did not respond for roughly forty years. The reason was that factories had been architected around a central steam shaft, and simply bolting an electric motor onto that architecture bought you almost nothing. The gains arrived only when engineers rebuilt the factory itself around what electricity made possible: a small motor on each machine, which meant machines could be arranged by workflow instead of by proximity to the shaft, which meant the assembly line. The technology was proven for decades before the transformation, because the transformation was never in the technology. It was in the rewiring.

Now look again at MIT's 5 percent. External partnerships, customized tools, deep integration into one specific workflow, twice the success rate of generic internal deployments. Look at where agents actually worked in 2025: coding, where output is verifiable by compilers and tests, where the loop is tight, where the workflow was rebuilt around the agent rather than the agent dropped into the old workflow. The pattern is exactly David's pattern. Agents succeeded precisely where someone did unit-drive work, redesigning the process around the new machine, and failed where organizations bolted a demo onto a steam-shaft workflow and waited for the output of the company to change materially.

That is the deepest sense in which 2025 was the year of the demo. In the electrification story, the demo era was not a failure. It was a stage, the one where the technology proves itself in showcases while the surrounding world has not yet been rebuilt to receive it. The mistake is not building demos. The mistake is booking the demo as if it were the deployment, forty years early.

What to do with this in 2026

Four questions and a habit, all of them things you can apply this quarter.

When a vendor, or your own team, shows you an agent demo, ask first: how many unattended steps is my real workflow, and how many was that demo? If the demo was 3 steps and your workflow is 20, the burden of proof is not partially met. It is untouched, because success does not scale linearly with steps; it decays exponentially.

Second: who catches the per-step errors in production? In the demo, the answer was the person driving it. If the production answer is "nobody," you are not deploying the demo, you are deploying its failure rate, compounded. The deployments that worked in 2025 had compilers, tests, review gates, or humans at the checkpoints.

Third: is the output verifiable at all? The agent wins of 2025 clustered where correctness could be checked cheaply, code that runs, citations that resolve. If nobody can cheaply say whether the agent's output was right, you will discover its error rate the way Klarna did, in customer satisfaction data, months later.

Fourth: who is doing the rewiring? MIT's 67-versus-33 split is the dynamo lesson priced in 2025 dollars. If the plan is to drop a general tool into an unchanged workflow, you have chosen the failing side of the split. Budget the workflow redesign, or do not budget the agent.

And the habit, which I mean as a genuine compliment to Altman: write the sentence. His prediction was dated, specific enough to score, and hedged where he was uncertain, which is exactly why we can have this conversation about it, and why he gets partial credit rather than a verdict of vibes. Most of the industry's 2025 claims cannot be scored at all, which is worse than being wrong. So write your own: "By January 2027, we will have agents doing X unattended, measured by Y." Put it somewhere you cannot quietly delete. The year of the agent taught us that demos are promissory notes, and the note comes due not when the model improves but when someone does the rewiring. Score your sentence in a year. That is how the demo era ends: not with a better demo, but with people who kept the receipts.


Sources

The demo era ends not with a better demo, but with people who kept the receipts.

The two questions this essay says to ask of any agent deployment — who catches the per-step errors, and is the output verifiable at all — are the questions the agent trust stack exists to answer. It ships as separate installable pieces so you wire in exactly the checkpoints the compounding arithmetic demands: verification against ground truth where correctness can be checked, a provenance record of what the agent actually did so you can score the work later rather than discover its error rate in the customer-satisfaction data months on, and ratings that price an agent by its real track record. Klarna found its effective error rate the slow way. The receipts are how you find it before it compounds.

Read the Theory of Agent Trust  ·  Hosted Chain of Consciousness

pip install agent-trust-stack  ·  npm install agent-trust-stack

Or the provenance record on its own: pip install chain-of-consciousness / npm install chain-of-consciousness.