← Back to blog

Stanford Says 12% to 66%, but 12% of What?

The same report holds two twelves that tell opposite stories. What separates a floor from a ceiling is never the number. It is the denominator, and the room the number was measured in.

Published July 2026 · 10 min read · AI benchmarks / evaluation / measurement literacy


The 2026 AI Index, Stanford's big annual audit of where artificial intelligence actually stands, landed this spring with a number that traveled fast. AI agents, the report said, had gone from succeeding at 12% of real computer tasks to 66.3% in a single year. Divide one figure by the other and you get the line that made the rounds: a 5.5-fold improvement in twelve months, agents now within 6 percentage points of human performance at operating a computer. It is a genuine finding from a serious source, and it is exactly the kind of number that ends up in a board deck by Thursday. So let me ask the only question that really matters about it, which happens to be the question in the title, and which is not rhetorical. 12% of what?

The “of what” is a benchmark called OSWorld: 369 specific tasks spread across real applications and operating systems, each one set up in a known starting state, with a script standing by to check, unambiguously, whether it got done. So “12% to 66%” means an agent went from finishing roughly 44 of those 369 tasks to finishing roughly 245 of them. That is a real and large jump, and I want to be clear at the outset that it is not fake and this essay is not going to pretend it is. But notice the word doing the quiet persuading: real. “Real computer tasks” is technically accurate, since they use real applications, and it invites you to picture your computer, your inbox, your genuinely messy desktop, when what was measured is a curated, gradeable set of 369. Even the human baseline it came “within 6 points” of is only about 72%, which is to say people fail more than a quarter of these tasks too. The number is honest. The picture it plants in your head is not the number.

And here is where the AI Index does something almost too neat, because a few pages away in the very same report sits another 12%, and it means the opposite thing. Robots, the report says, succeed at only 12% of real household tasks. On BEHAVIOR-1K, a benchmark built around a thousand everyday household activities, the best team managed full success about one time in eight. And the same report notes that in simulation, on a benchmark called RLBench, robotic manipulation reaches 89.4%. Sit with that pairing for a second. Two twelves, in one document, telling opposite stories. The agent's 12% was a floor, a starting line the curve sprinted away from on its way to 66. The robot's 12% is a ceiling, the hard limit of what actually works once you take a machine that scores 89% in a simulator and set it loose in a real kitchen. Same digit. One is the bottom of a ramp; the other is the top of a wall. What separates them is not the number. It is the denominator, and the room the number was measured in.

That is the entire reflex the title is trying to install, and the robots make it visible in a way the software agents cannot manage on their own. The gap between “89.4% in a simulation” and “12% in a real house” has a name roboticists have been fighting for decades: the sim-to-real gap. A system can be nearly flawless in the controlled, repeatable world of a simulator and nearly helpless in the uncontrolled, once-only world of an actual room, because the simulator quietly deleted everything that made the real task hard: the friction and slippage of physical objects, the lighting that is never twice the same, the cup that sits an inch to the left of where the training data left it, and the small brutal fact that you get exactly one attempt and cannot reset the kitchen. Now look back at the software agents with that fresh in mind. OSWorld is their simulator. It is a cleaner, more repeatable, more gradeable world than the one your actual workflow lives in, and the 66.3% was earned inside it. The distance between “66% on OSWorld” and “runs your operations unattended” is the same shape of distance as the one between “89% on RLBench” and “does your laundry.” It is a sim-to-real gap wearing software clothes. The digital agents are not exempt from the wall the robots are stuck against. They are simply earlier on the same curve, and their simulation happens to have a friendlier name.

There is a second thing the number hides, and once again the report hands it to you a few lines later. 66.3% is an average, and averages over AI capability are unusually treacherous, because the surface underneath them is what the report itself calls jagged. The same 2026 AI Index that reports 66% on computer tasks also reports that Google's Gemini Deep Think won a gold medal at the 2025 International Mathematical Olympiad, scoring 35 points while working entirely in natural language, inside the same 4.5-hour limit the human competitors get, up from a silver the year before. And the same report notes that the best model reads an analog clock correctly only about half the time, against roughly 90% for humans. A machine that can take gold at one of the hardest mathematics competitions on Earth cannot reliably tell you it is a quarter past three. That is not a small blemish on an otherwise smooth ability. It is the texture of the whole thing. So when a single figure says 66%, it is reporting the mean over the basket and telling you nothing about which third it fails, and those failures are not sprinkled evenly. They clump on exactly the things a benchmark is worst at capturing: perception, long horizons, the un-templated task nobody thought to write a grader for. And that clustering is not bad luck, it is structural. A benchmark can only score what it can automatically check, so it fills up with the checkable and leaves out the rest, which means the parts of your work that resist a clean grader are precisely the parts that never make it into the basket and therefore never appear in the 66. The report's own summary admits as much, conceding that the agents still fail roughly one attempt in three. And the third they fail is, with grim reliability, the third your real work is made of.

You can watch the same trick play out on a second benchmark in the same report. On SWE-bench, where models are handed real bug reports and asked to fix them in real codebases, scores climbed from 60% to nearly 100% in a single year, a number that sounds like software engineering just got solved. But the benchmark being quoted is SWE-bench Verified, and the word Verified is doing exactly the work the word real did for OSWorld. It is a human-screened subset, filtered down to the problems annotators judged well-specified and reliably gradeable. So “near 100%” means near 100% of the issues that were pre-selected to be winnable, not near 100% of the bugs sitting in your repository this afternoon. The denominator is hiding in plain sight, inside the benchmark's own name, and almost nobody who forwards the number stops to go looking for it.

One more caveat, and I will keep it short because a companion piece on this blog takes it apart in full. OSWorld itself changed during the very year the number climbed. In July 2025 its authors shipped a repaired release, OSWorld-Verified, that fixed broken tasks and faulty graders. If the 12% floor was measured on the old version and the 66.3% on the cleaned-up one, then part of the 5.5-fold gain is the ruler getting more accurate rather than the runner getting faster. That does not erase the progress, but it does mean the two endpoints may not be the same 369 tasks, which is one more reason to treat “5.5 times in a year” as a claim to inspect rather than a fact to forward. That companion piece works through what happened when one model, GPT-5.4, crossed the human line on this same benchmark, and why the people who build agent systems did not change a single thing in response. This piece is about the number itself.

So here is the practical residue, the thing you can actually use the next time a capability number comes flying past you, which will be soon. Three questions, in order. First: of what basket? A percentage means nothing until you know the set it is a percentage of, and “real” in a benchmark's name is a marketing word, not a measurement. Second: how jagged inside? An average hides its own worst regions, and for AI the worst regions can be genuinely startling, clocks rather than calculus. Third, and this is the one that decides whether the number has anything to do with you at all: is your task in the basket, or in the failing third? Run a real claim through it and watch how fast it changes shape. A vendor tells you their agent books travel with 90% reliability. Of what basket: a tidy set of demo itineraries, or the edge-case-ridden reality of actual corporate travel? How jagged inside: does the 90% hold on the international multi-leg trip with a visa and a schedule change, or only on the round-trip nobody needed help booking? And is your use in the basket at all, or is it the exception that made you want an agent in the first place, the exact case a demo would never include? Ten seconds of those three questions and the 90% either firms up into something you can plan around or quietly evaporates. The 66% was earned on templated, checkable tasks in a friendly environment. If your work looks like that, the number is good news and you should act on it. If your work is the open-ended, once-only, nobody-graded-it kind, which is what most real work is, then you are standing on the far side of a sim-to-real gap the headline never mentioned, and the honest figure for your situation is somewhere well south of 66 and impossible to read off a press release.

Notice that the same reflex just handled three different numbers in three different domains. A computer agent's leap to 66, a household robot stuck at 12, a coding model's climb to near 100, and in every case the move was identical: find the basket the percentage is a percentage of, feel around inside it for the jagged edges, and ask whether your actual problem lives in the part that got measured or the part that quietly did not. It is a single skill, and it does not care in the slightest what the benchmark is named or which lab is doing the announcing.

The improvement is real, and that has to be said as plainly as the skepticism, because the opposite mistake, waving away a genuine 5.5-fold gain as pure hype, is exactly as lazy as swallowing it whole. Agents really did get dramatically better at a large, curated set of computer tasks in one year, and even setting the screened subset aside, the coding models really did make a large and verifiable jump on the problems they were handed. Something is happening, and it is fast. But the two twelves in that report are the lesson worth carrying out, because between them they prove that the identical number can be a floor or a ceiling depending entirely on what it counts and where it was counted. Read the number. Then read the room it was measured in, because the room is where the number either survives contact with your actual problem or quietly comes apart in your hands. Stanford told you 12 to 66. The durable skill is never taking the of-what for granted again.


Sources

A benchmark number is a percentage of a basket you did not choose. Your work lives in the part that did not get measured.

The three questions, of-what-basket, how-jagged, is-your-task-in-it, are exactly the questions you cannot answer from a headline score, and they are the questions that decide whether an agent is safe to hand your real work to. The antidote to a decontextualized number is grounding: a record of what an agent actually did on your tasks, a reputation earned on real work rather than a curated set, and verification that survives contact with the once-only case. The Agent Trust Stack is that grounding, provenance, reputation, and verification, so an agent's fitness for your problem is something you can check instead of a percentage you have to take on faith.

Read the Theory of Agent Trust

Install the whole stack: pip install agent-trust-stack  ·  npm install agent-trust-stack