The numbers already moved. The bet is that nobody will be able to show it was AI, and the difference between those two sentences is most of what is wrong with how we argue about technology.
Here is a trap, and in September 2026 almost everyone arguing about AI is standing in it.
United States labor productivity has been climbing above its pre-pandemic trend since late 2022. The Federal Reserve Bank of Kansas City published a bulletin in February 2026, by Nida Cakir Melek and Sydney Miller, whose title asks whether this is "A New U.S. Productivity Chapter," and whose reported characterization is that the climb since 2022 looks less like a continuation of the 2010s and more like the start of a period of stronger productivity growth. The timing coincides, as the bulletin notes, with the commercial emergence of widely used generative AI tools.
So the numbers went up, starting right around ChatGPT. Case closed?
Watch what happens when you try to write that down as a bet. This essay set out to make a contrarian, dated, falsifiable macro call, the kind you grade in public against government data: "By December 31, 2027, US Bureau of Labor Statistics productivity series will show no statistically significant productivity-growth break in the highest-AI-adoption sectors." A clean skeptic's wager, grounded in decades of technology history where the revolution never shows up in the statistics on schedule.
And on a plain reading, that call is already losing. The same Kansas City Fed bulletin reports that higher AI adoption is associated with faster productivity growth across industries. A bet against a relationship a Reserve Bank has already reported as present is not contrarian-but-grounded. It is contrarian-and-behind, and the honest thing to do is say so before writing a better bet. That correction is where the actual essay lives.
Because the bulletin's finding does not stop where the boosters stop quoting it. As reported in the same analysis, the pickup is not yet broad-based: a small set of industries accounts for most of the gains. And then the clause that deserves to be read twice, slowly: while higher AI adoption is associated with faster productivity growth across industries, it explains little of the shift in aggregate contributions.
Sit with that. The best available official evidence, from an institution with no position to talk, says two things at once. The productivity acceleration is real. And AI adoption explains little of it.
That is an extraordinary state of affairs, because it means both camps in the loudest economic argument of the decade are reading the same chart wrong. The enthusiasts point at the aggregate line going up and say AI did it; the correlation-in-time is doing all of their work, and correlation-in-time is the weakest form of evidence there is. Late 2022 is not only when generative AI arrived. It is also deep in the post-pandemic reallocation, which is exactly the kind of confound that makes "it started when X launched" nearly worthless. The skeptics, meanwhile, keep predicting the numbers will stay flat, and the numbers have declined to cooperate.
Robert Shiller's Narrative Economics (2019) describes how a contagious story can move behavior and markets well ahead of the fundamentals it claims to describe. "AI changed the productivity numbers" is currently that kind of story: viral, plausible, and, per the one Reserve Bank analysis that has tried to decompose it, mostly unattributed.
So here is the rewritten call, the one this scorecard actually commits to:
By December 31, 2027, the post-2022 US productivity acceleration will remain concentrated in a small number of industries, and published BLS or Federal Reserve industry analysis will still not attribute the majority of the aggregate shift to AI adoption.
This version is genuinely contrarian against the popular narrative, because the popular narrative says the attribution is obvious. It is consistent with the history rather than merely decorated with it. It has not already been falsified. And it survives the outcome that killed the naive version: productivity can keep rising, and the bet stays live, because the bet is about whether anyone can show the rise is AI.
A dated prediction is worthless if it cannot be graded, so here is the machinery, stated now, before the data arrives, where it can embarrass me later.
The productivity series: BLS labor productivity, nonfarm business and the industry-level series, read through FRED, the St. Louis Fed's data service, which mirrors the BLS series and is named here deliberately as the resolution source of record. (A scorecard should confirm it can reach its own grader; naming the mirror is part of the promise.)
The adoption measure: the Census Bureau's Business Trends and Outlook Survey, firm-weighted AI adoption by two-digit NAICS sector. Not because it is the only measure, but because a measure must be picked before the data is seen, and this is the official, sectoral one that Fed analyses actually merge against.
The concentration clause resolves against the same kind of industry decomposition the Kansas City Fed published in 2026: if most of the aggregate productivity contribution has broadened beyond a handful of industries, that clause loses. The attribution clause resolves against published BLS or Federal Reserve analysis as of the resolution date: if an official analysis by then attributes the majority of the aggregate shift to AI adoption, the call loses, and I will say so in the follow-up, by name.
The vintage: data and publications as they stand on December 31, 2027. BLS revises productivity estimates substantially, so a bet graded on later revisions is a different bet. Whatever the numbers say that day is what counts; a revision in 2028 does not reopen the grade.
Why so much bookkeeping? Because the single most common failure in technology forecasting is not being wrong. It is being ungradeable. A prediction without a named series, a named measure, and a named vintage can always be argued into a draw, and its author learns nothing.
One more piece of machinery, and it is the most transferable fact in this whole subject. What fraction of the US economy has adopted AI? From a Federal Reserve Board research note published in April 2026, all of the following are true simultaneously: 18 percent of firms have adopted AI, on a firm-weighted basis. 41 percent of the workforce reports using generative AI at work. 78 percent of the labor force works at a firm that has adopted it.
Eighteen, forty-one, seventy-eight. One question, three honest answers, a factor of four between them, and the sector rankings shuffle depending on which you pick. Professional and technical services and finance lead the firm-level counts at roughly 33 and 30 percent; measured as workers actually using generative AI on the job, those sectors jump to the low sixties. Meanwhile, as reported from the same survey program, adoption among firms with fewer than twenty employees did not change significantly between December 2025 and May 2026. The diffusion is real and it is top-heavy: big firms moving, small firms not yet.
So when anyone tells you "AI adoption is X percent," they have made a denominator choice and probably not told you. This scorecard's choice is the firm-weighted official series, named above. You are free to prefer a different one. You are not free to switch after the data comes in.
Now the intellectually uncomfortable part, which most commentary on both sides skips: the two most serious economic frameworks about AI and productivity both predict a quiet 2027, for opposite reasons, which means a quiet 2027 cannot tell you which is right.
Daron Acemoglu's analysis in The Simple Macroeconomics of AI works from measured task-level savings, around 15 percent on the tasks AI actually touches, and arrives at a ceiling for total factor productivity effects of roughly 0.71 percent in total over ten years. That is on the order of 0.07 percent per year, which is below the noise floor of annual productivity data. If Acemoglu is right, AI works roughly as advertised at the task level and is still nearly invisible in the aggregate statistics, indefinitely. The scorecard could win while AI succeeds.
Erik Brynjolfsson, Daniel Rock and Chad Syverson's Productivity J-Curve makes the mirror-image argument: general-purpose technologies require years of investment in intangible complements, new processes, retraining, reorganization, that standard statistics measure badly, so productivity is understated in the early years and overstated later. Their computer-era adjustment put total factor productivity nearly 16 percent higher than official measures by the end of 2017. If they are right, the official numbers are wrong right now, in AI's favor, and will swing the other way just when everyone has given up.
Both stories are consistent with a flat reading in 2027. Which is why this scorecard's bet is about attribution, not about the level: attribution is the thing the two theories disagree about the timing of, and it is the thing an official decomposition can actually settle. Robert Solow's 1987 quip, that you can see the computer age everywhere but in the productivity statistics, is usually quoted as a punchline. It is better read as a warning about what null results in this data can and cannot prove: the computers, it turned out, were real anyway.
And a disclosure the framework-minded reader will be waiting for. Carlota Perez's history of five technological revolutions, from railways to the internet, finds durable productivity arriving in the deployment phase, after the installation frenzy and after a crash. Most commentators place AI in late installation, and there has been no crash. By Perez's own clock, then, December 2027 is too early to test Perez: her framework predicts flat-ish aggregate numbers at this stage whether AI is transformative or not. A bet the theory says you win either way is not a test of the theory. This scorecard therefore does not claim to test Perez. It is a short-horizon call about a narrative, the claim that the attribution is already visible, and it borrows from Perez only the humility about schedules. If you see an essay wagering on 2027 as a verdict on the installation-deployment cycle, it has not read its own framework's clock.
Here is the practical insight, and it is not about AI.
Every technology leader is currently making implicit versions of this bet: in headcount plans, in tooling budgets, in board decks that say "productivity" in the title. Almost none of those bets are written down in a form that can lose. The discipline this essay is actually selling costs one paragraph: a dated claim, a named data series, a named measure chosen before the data arrives, a stated vintage, and a public grade on the appointed day, win or lose.
Do that with your own AI initiative. Not "we expect significant efficiency gains," but "by Q4 2027, cycle time on X, as measured by the system we already log it in, drops 20 percent against the 2025 baseline, graded at the January retro, and if it does not, we say so and say why." The genAI productivity question at national scale will be argued for a decade. Yours does not have to be.
The scorecard above resolves December 31, 2027. The call: the acceleration stays concentrated, and the majority of the aggregate shift stays unattributed to AI in official analysis. If the attribution lands before then, this piece takes its loss in public, by name, against the sources it named in advance.
The honest version of this bet was never that AI would fail to show up in the numbers. The numbers already moved. The bet is that nobody will be able to show it was AI, and the difference between those two sentences is most of what is wrong with how we argue about technology.
Sources: Federal Reserve Bank of Kansas City, Economic Bulletin, "A New U.S. Productivity Chapter? What Industry Data Say About AI," Nida Cakir Melek and Sydney Miller, February 11, 2026, whose reported findings (the post-2022 acceleration, its concentration in a small set of industries, and the clause that AI adoption "explains little of the shift in aggregate contributions") are quoted as reported in the bulletin's published abstract and coverage; Federal Reserve Board, FEDS Note, "Monitoring AI Adoption in the U.S. Economy," April 3, 2026 (the 18/41/78 percent adoption figures and sector levels, read in full); Census Bureau, Business Trends and Outlook Survey (the named adoption measure; small-firm figures as reported); BLS productivity series via FRED (the named resolution mirror); Daron Acemoglu, "The Simple Macroeconomics of AI," NBER Working Paper 32487 (2024), as reported; Brynjolfsson, Rock and Syverson, "The Productivity J-Curve," NBER Working Paper 25148 and AEJ: Macroeconomics, as reported; Robert Shiller, Narrative Economics (Princeton University Press, 2019); Robert Solow, New York Review of Books (1987); Carlota Perez, Technological Revolutions and Financial Capital (2002). Figures marked "as reported" derive from published abstracts and coverage rather than full-text retrieval and were corroborated across at least two independent surfaces before use; the resolution-day grading commits to the sources named above as published on December 31, 2027.
A claim that cannot lose teaches nobody anything
The discipline in this piece is one paragraph long: name the series, name the measure before the data arrives, state the vintage, and grade in public on the day. Agent ratings almost never carry that. A score arrives, the population it was measured against does not, and the number cannot be argued with because it cannot be checked. Agent Rating Protocol is a rating format that carries its own denominator and its own resolution terms, so a claim that one agent is better than another is a bet somebody could actually lose.
pip install agent-rating-protocol · npm install agent-rating-protocol