A self-reported accuracy score is a fact about the grader, not about the future.
In October 2010, Ray Kurzweil published a 148-page document called "How My Predictions Are Faring." In it, he went back through the 147 predictions he had made for the year 2009 in his 1999 book The Age of Spiritual Machines, and graded them, one by one, himself. The tally: 115 "entirely correct," 12 "essentially correct," 17 "partially correct," 3 wrong.
Add the first two buckets and you get the number that has followed him ever since: 127 of 147, 86 percent.
That figure gets cited constantly, usually as evidence that the most famous futurist alive has a verified track record. Here's the thing, though: nothing about it is verified. It's a self-assessment, scored by the author, against a rubric the author invented, using interpretations the author chose. And it turns out that when other people grade the same 147 predictions, they get wildly different numbers: 50 percent, 25 percent, depending on who holds the red pen.
The gap between 86 and 25 is not a disagreement about what happened in 2009. Everyone agrees on what happened in 2009. The gap is a measurement of something else entirely, and it's the thing this essay is actually about: a self-reported accuracy score is a fact about the grader, not the future.
Start with the rubric. "Entirely correct" is a defensible category, since a prediction either described the world or it didn't. But look at the second bucket, the one doing the work of lifting 78 percent to 86.
Kurzweil defines "essentially correct" right there in the document: these are predictions that haven't come true yet but are "only a few years away," which he counts as hits "given that these predictions were specified by decades (not years)." Sit with that for a second. A prediction for 2009 that was false in 2009 gets scored as essentially correct because it might be true by 2012, and because a decade is a fuzzy thing. The pass bar moved after the exam.
The individual grades are where the rubric really flexes. Take his 1999 prediction that "most portable computers will not have keyboards." In 2009, every laptop had a keyboard, which looks like a clean miss. But in his self-assessment, Kurzweil counts MP3 players, cameras, phones, tablets, and game players as portable computers, at which point the majority of "portable computers" indeed have no keyboard, and the prediction is graded true. He states the interpretive principle openly: "It is not appropriate to use today's marketing categories to interpret the meaning of these earlier predictions."
Maybe that's fair. An iPhone genuinely is a computer. But notice what's happening: the prediction's meaning is being chosen after the fact, by the person being graded, from among the readings available, and the chosen reading is, reliably, the one that scores. A prediction that "personal computers are commonly embedded in clothing and jewelry" is graded via iPod nanos worn as pins and health monitors woven into garments. Each individual call is arguable. That's precisely the problem. When every ambiguous call is arguable, the grader's generosity becomes the dominant term in the final score.
In 2012, Stuart Armstrong, then a researcher at Oxford's Future of Humanity Institute, did the simplest useful thing anyone can do with a contested scorecard: he re-graded it.
His first finding arrived before he'd assessed a single prediction. Reading the same 2009 chapter of the same book, Armstrong counted 63 predictions where Kurzweil counted 147. Even the denominator is a grading choice, because it depends on how you split compound sentences into claims. Two honest readers of one chapter can't agree on how many predictions it contains, and you're supposed to trust a decimal-precision accuracy score built on top of that?
Armstrong then sampled ten predictions at random and scored them: five at least weakly true, four at least weakly false, one unclassifiable, so roughly 50 percent. From there he ran the numbers backward: if the true hit rate is what his sample suggests, how likely is it that a fair grading of all of them lands at 86? His answer: Kurzweil's self-score sits at the 99th percentile of the plausible distribution, the 94th even with generous adjustments. The title of his write-up is the fairest one-line verdict anyone has issued on the whole affair: "Kurzweil's predictions: good accuracy, poor self-calibration." His twin conclusions: "He's most likely good at predicting" and "He's most likely overconfident, reluctant to admit his misses."
The same year, Knapp at Forbes ran the adversarial version, taking twelve headline predictions and grading them strictly: one completely true, four partially true at half credit each, which he totals as "a score of 3 / 12 - or 25% accurate."
So the identical set of sentences about 2009 scores 86 percent when Kurzweil grades, ~50 percent when a neutral academic samples it, and 25 percent when a skeptical journalist picks the headliners. And before you conclude the harsh number is the "real" one, no. Knapp's 25 is also a grading artifact: a strict rubric applied to a hand-picked dozen. That's the uncomfortable landing point: on predictions written this loosely, no single scorecard is trustworthy, including the hostile one. The spread between the scorecards is the only honest measurement in the whole exercise, and what it measures is the graders.
And the pattern replicated out of sample. In 2020, Armstrong ran the exercise again on Kurzweil's predictions for 2019, this time properly: 105 separate statements, 34 volunteer assessors, 3,078 individual judgments. Headline result: "Kurzweil's predictions for 2019 were considerably worse than those for 2009, with more than half strongly wrong." The self-scorecard says 86 percent; the first crowd-graded scorecard says most-strongly-wrong. Both numbers are about the same man's foresight. They cannot both be measurements of it.
Here's the mechanism, stated plainly. A grade has three free parameters: what counts as a prediction (the denominator, 147 or 63 or 12), what counts as passing (entirely / essentially / partially correct), and what the words mean (is a camera a portable computer?). An accuracy score is only as objective as the least objective of the three.
Self-grading loose predictions surrenders all three parameters to one interested party, after the outcome is known. It isn't lying; nothing in Kurzweil's 148 pages is factually false, and that's what makes it instructive. It's that a vaguely worded prediction is a coin that the grader gets to call after the flip, and a self-grader calls every flip in their own favor, sincerely, one defensible interpretation at a time. The re-graders converged on this exact diagnosis: most of the predictions simply aren't worded precisely enough to evaluate objectively. The 86 percent wasn't manufactured in 2010 when he graded generously. It was manufactured in 1999, when the predictions were written softly enough to be gradable generously later.
If you want a general law for reading anyone's track record, it's this: partial credit is where track records go to launder. Whenever a forecaster reports their own accuracy, find the fuzzy bucket ("essentially correct," "directionally right," "true in spirit") and check how much of the headline number lives there. Here, eight points of the 86 live in "essentially correct," and another 12-percent tier of "partially correct" sits one interpretive nudge below the pass bar, ready.
If the essay stopped here, it would be a cheap shot, and it would miss the more interesting half of the story. Armstrong's verdict had two clauses, and the first one was "he's most likely good at predicting."
Because alongside the 147 soft predictions, Kurzweil made a handful of hard ones, and the hard ones are the ones a generous grader can't help. He has said "a computer will pass the Turing test by 2029" since 1999, and he set 2045 for the Singularity in The Singularity Is Near in 2005, and he has not moved either date in the decades since, through every AI winter headline and every wave of ridicule. In 1999, "human-level AI by 2029" was a punchline in serious company.
Now look at it from 2026. In 80,000 Hours' review of expert AGI forecasts, the Metaculus forecaster community averages a 25 percent chance of AGI by 2029 and 50 percent by 2033, as of February 2026, down from a median of fifty years away as recently as 2020. That last clause is the number worth holding onto: not where informed forecasters sit today, but how far they have moved, and in which direction. Whatever you believe about the destination, the sober middle of opinion has migrated toward Kurzweil's date, not away from it. His single most falsifiable prediction, a date, on a public claim, held for 27 years, is the one part of his record that currently looks strong without anyone's generosity.
That's the real tragedy of the 86 percent, and the reason "he's just a hype man" is the wrong takeaway. The self-graded number didn't protect his reputation; it spent it. Every "86% accurate" citation invites the Forbes re-grade, and the argument that follows drowns the genuinely impressive thing, which is that a compute-extrapolation argument from the 1990s put a date on machine intelligence that the forecasting community of 2026 has been steadily moving toward. A good forecaster with a self-serving scorecard ends up less credible than his own forecasts deserve. The generosity doesn't invent a signal; it contaminates one. For anyone keeping a professional track record, that's the sharpest lesson here: inflation isn't just dishonest, it's expensive. You pay for it out of your true wins.
Which brings us to what Kurzweil is saying right now, and a pattern worth naming gently. On Peter Diamandis's Moonshots podcast in early 2026, he reaffirmed the date and then kept talking: "In 1999, I predicted 2029 for AGI and I still predict 2029. Elon Musk says 2026. I think we'll have a lot of things that remind us of AGI, but we really won't be convinced in 2026. Maybe 2027, 2028. By 2029, I think everyone will accept it."
Read that second half twice. "Things that remind us of AGI." "Won't be convinced until." These are not resolution criteria; they're the opposite, a description of arrival so experiential that any future can be graded either way. The man who graded his 2009 predictions with an "essentially correct" bucket in hindsight is now pre-installing the same bucket in foresight, before 2029 arrives to be graded. To be fair, he's describing something real. 2026 genuinely is full of systems that "remind us of AGI," and reasonable people disagree daily about what would settle the question. But that's exactly why the phrasing matters. "AGI by 2029" is a date attached to a word whose meaning is dissolving as it approaches. Unless the claim gets a resolution rule, a specific test, named in advance, graded by someone else, the sharpest prediction in futurism is being quietly converted back into the soft kind, the kind its author has never once graded himself as having missed.
The fix is not "be more skeptical of Kurzweil." The fix is structural, it's old, and every field that grades forecasts for a living already uses it. Weather forecasting solved this decades ago: you say 70 percent chance of rain, the rain falls or doesn't, and a Brier score does the grading, with no essentially-correct bucket and no post-hoc reading of what "rain" meant. Same discipline, three moves:
We try to practice this in public ourselves. Three Dated Bets on Agentic-AI Insurance puts three claims about the agent-insurance market on the record against a single resolution date, January 1, 2027, written before the outcome and gradable by anyone who cares to check on the day. Not because our forecasts matter more than anyone else's, but because the alternative is ending up with our own 86 percent: a number that sounds like a track record and functions as a confession that nobody else was allowed to check.
And that's the way to watch 2029 arrive. One of the most consequential predictions ever made comes due in three years, made by a man who is genuinely good at this and demonstrably generous to himself about it. We will find out whether he was right. We just won't find out from him, because the only scoreboard that settles anything is one the forecaster doesn't control. Kurzweil's own best prediction, the dated one, is the single part of his legacy that never needed his thumb on the scale. The kindest thing history can do for the 86% prophet is grade him strictly.
Any score is a claim about who graded it, against which rubric, with what left open.
The 86 percent was not a lie. It was a rating whose grader, rubric and denominator all belonged to the party being rated, and nobody could see that from the number alone. The same hazard applies to every agent rating you will ever read. Agent Rating Protocol makes the grading structure explicit: what was scored, on which criteria, by whom, so a score can be interrogated rather than taken at face value. A rating you cannot re-grade is a self-assessment wearing a decimal point.
Read the Theory of Agent Trust
pip install agent-rating-protocol · npm install agent-rating-protocol
Or the whole trust stack at once: pip install agent-trust-stack / npm install agent-trust-stack