Here is the story everyone tells, and the story this essay was originally assigned to tell: in 2022 and 2023, Gary Marcus and his fellow skeptics predicted an AI winter: fundamental limitations of large language models, hype outrunning reality, an implied pullback. Then the pullback never came, capex went vertical, and the skeptics were buried by the market they doubted.
It is a satisfying story with one problem. The founding document doesn't contain the prediction.
Go read the actual March 2022 essay, "Deep Learning Is Hitting a Wall," in Nautilus. The word "winter" does not appear in the text. There is no forecast of investment collapse, no revenue decline, no funding pullback. What the essay actually argues, in its own words: "Making GPT-3-like models bigger makes them more fluent, but no more trustworthy." And: "In time we will see that deep learning was only a tiny part of what we need to build if we're ever going to get trustworthy AI." The conclusion is not "sell"; it is "hybrid AI, not deep learning alone seems the best way forward." The closest thing to an economic claim is a warning that tens of billions invested in scaling specifically could turn out to be for naught. That is a claim about a route, not about a market.
Marcus has been saying this about his own essay for years, with the weary precision of a man correcting the same misquote at every dinner party: the 2022 piece "was neither about revenue or AI's potential upper limits" but "an argument that the pure scaling of LLMs would not get us to AGI." You don't have to take his later summary on faith; that is why I quoted the 2022 text directly. The prophet of the AI winter never prophesied one. The winter was projected onto a technical critique by readers who needed the argument to be about money, because money was the argument they knew how to have.
Which sets up a much stranger question than "was the skeptic wrong?" Namely: what actually happened to each side's real predictions? Because the answer is that both camps were substantially right — about different propositions — and the four-year shouting match was powered almost entirely by nobody noticing they weren't disagreeing.
The market first, so there's no suspicion this is a skeptic's essay wearing a referee shirt.
If a winter means capital retreat, the record is not ambiguous; it is a rout. The five biggest infrastructure spenders (Microsoft, Google, Amazon, Meta, Oracle) went from roughly $238 billion of capex in 2024 to roughly $429 billion in 2025, a 73 percent jump, with 2026 plans in the $660–690 billion range per analyst tallies. Widen the lens to the fourteen largest data-center operators and the 2026 plans run near $750 billion against about $450 billion the year before. BloombergNEF counted more than 23 gigawatts of data-center capacity under construction at the end of September 2025, about three-quarters of it in the United States. Gartner put worldwide AI spending around $1.5 trillion in 2025 and above $2 trillion for 2026. (Caveat where it belongs: these are analyst and trade aggregations, and the 2026 figures are plans, not actuals. Plans can be cut. But nobody cuts from these altitudes into a winter.)
The single most diagnostic line in that record, though, is not a number. It is the hyperscalers' own repeated characterization of their situation: supply-constrained, not demand-constrained. Sit with that phrase, because it does real work. An AI winter, the historical kind of 1974 or 1987, is a demand collapse: the buyers stop believing, the money leaves, the field contracts. "We cannot build capacity fast enough for the orders we have" is not a mild version of that condition. It is the opposite condition, by definition. Whatever is happening, winter is the one word the evidence rules out.
Now the other scoreboard, and it deserves the same directness.
Grade the 2022 essay against 2025–2026 on the claims it actually made. Pure scaling reaching AGI: did not happen; nothing broadly recognized as AGI exists, and GPT-5's August 2025 launch landed to famously muted reception, the flagship data point that "bigger" had stopped translating into "transformative." Reliability solved by scale: hallucinations remain unsolved, four model generations after the essay said scale wouldn't solve them. Agents, the year's marquee product category, were, by year-end reviews, hyped and not reliable. The essay's central technical sentence, more fluent but no more trustworthy, reads today less like a hot take than like a product review written three years early.
One retrospective tallied sixteen of seventeen of Marcus's "high confidence" 2025 predictions as correct. Handle that number with tongs: "high confidence" is a self-selected subset (the forecaster's own safest calls), so 16/17 on that slice is not 16/17 overall. There is a neutral scoreboard, a third-party tracker called Forecast Fools that grades nineteen of his predictions; I tried to read it and it refused the connection, so this essay declines to print an overall accuracy percentage it couldn't verify. But the individual items above (no AGI, muted GPT-5, unreliable agents, unsolved hallucinations, almost nobody but Nvidia profitable on AI) are corroborated separately, and they are enough for the honest headline: the most famous AI skeptic has a strong recent forecasting record on the claims he actually staked.
And fairness cuts both ways, so grade the optimists on their real claims too. The 2022 reaction to the wall essay was swift and vitriolic: thousands of people ridiculed it, and Sam Altman took public shots within weeks. Yann LeCun was still at it in late 2024: "Deep Learning is absolutely not hitting a wall. He was wrong when he first said it, and he is still wrong about that today." On the narrow claim that deep learning remains the foundation of every system that matters, LeCun is right. On the claim implied by the ridicule — that scaling would carry through to general intelligence and the reliability problems would dissolve on the way — the scoreboard reads muted GPT-5, persistent hallucination, agents that demo better than they deploy. Both camps' confident wings missed. The difference is which miss got a season named after it.
So how did "scaling won't reach AGI" become "winter is coming" in the public record? Through two words that each quietly meant two things.
"Wall": Marcus's own title, and the title is the part of the essay that was wrong, in the sense that titles are claims and this one overclaimed. Deep learning did not stop; it kept improving, kept absorbing capital, kept shipping. What hit a wall, on the evidence, was a specific conversion rate: parameters into trustworthiness, benchmarks into deployment. "Wall" let supporters hear "progress stops" and forced critics to defend a stoppage nobody could see.
"Winter": the other side's word, and the sharper equivocation. A technical ceiling on one route and a market contraction are different events with different observables, and "winter" fused them, so that every quarter of record capex could be scored as a refutation of a claim about cognition. Marcus forecast a technical ceiling on a route. The market forecast returns. Both could be right at once, and, on the record so far, both were. The winter never came because a winter was never what the skeptical case, as written, predicted.
For anyone who runs on arguments for a living, this is the transferable defect: the debate persisted for four years because the proposition was never pinned. Neither side could lose, so neither side could update.
Which brings us to the actual state of affairs in early 2026, and it is stranger than either camp's ending.
Capital is arriving at 60–75 percent annual growth, into supply-constrained markets. Simultaneously (and this is the commentary consensus, flagged honestly as assertion rather than statistic, because a real graduation-rate number does not seem to exist in public) most enterprise AI in 2026 still sits in the bucket of pilots that do not graduate to production. Reliability, the exact axis the 2022 essay named, is the most commonly cited reason.
A winter is legible: money leaves, progress stops, everyone agrees on the diagnosis. This is something else: investment and deployment running on separate clocks, record spending stacked on top of an unclosed reliability gap, the two curves diverging for years without forcing a reconciliation. A bull reads the capex as proof the gap must be closing; a bear reads the gap as proof the capex must be wrong. The uncomfortable both-and reading, the one the evidence actually supports, is that the market is pricing the eventual closing of a gap that the technical critique correctly says is not yet closed by scale alone. That is not a winter or a vindication. It is an open bet, at the size of the capex figures above, on which side of the equivocation resolves first.
The honest thing to do with that is not to resolve it. It is to say what would.
Not a capex number. Capex measures conviction, and conviction is what both manias and infrastructure buildouts look like from the outside; the same $429 billion is compatible with either story. What would settle it is a graduation rate: of the pilots started, what fraction reached production and stayed there? That number would adjudicate the actual disagreement — does the reliability gap close fast enough to justify the buildout — and it is conspicuous that in a field drowning in benchmarks, nobody publishes it. If you want to know which way this breaks before the market tells you, the graduation rate inside your own organization is the leading indicator you already own.
The portable lessons, then.
Recover the proposition before you grade the prophet. The received story about Marcus was checkable in one hour with the primary text, and it was backwards. Most "X predicted Y and failed" stories in your field are compressions by people who read the title. Before you update on a failed prediction, or mock one, find the sentence where the prediction was actually made. Titles are what get remembered, and titles are where authors overclaim.
Name your claim's observable, or you've made two claims. "Hitting a wall" and "winter is coming" both smuggled a second proposition inside a vivid phrase, and each side spent years refuting the half the other never staked. When you forecast (a technology, a migration, a competitor) state what would count as wrong in a unit someone else can check: a route, a rate, a date. An argument that can't be lost can't be won, and it also can't end.
And distinguish demand problems from conversion problems. The one-question diagnostic that cuts through every "is the boom real" argument: is the constraint supply or demand? Then the follow-up that cuts through the answer: fine, demand is real, but what fraction of it converts to production? A business can be genuinely supply-constrained on pilots and still starving on deployments. That is not a winter. It is the specific, stranger weather we are actually in, and it doesn't have a season name yet.
Sourcing notes: the 2022 essay's key sentences are quoted from the Nautilus text directly, which was read for this piece; the finding that "winter" does not appear in it and that it contains no investment-collapse forecast is from that read. The Forecast Fools tracker (the neutral scoreboard for Marcus's predictions) returned HTTP 403 on access, so no overall accuracy percentage is printed; the 16-of-17 figure is characterized as a sympathetic-source tally on a self-selected subset and load is carried instead by individually corroborated items. Capex figures are analyst/trade aggregations, with 2026 numbers labeled as plans; the "pilots do not graduate" claim is labeled as commentary-grade assertion because no public graduation-rate statistic was found. The LeCun quote is from his public Threads post; the description of the 2022 reaction (ridicule, Altman's shots) is as reported in the cited accounts.
The essay's own diagnostic is that an argument without a checkable unit cannot end. The same is true of an agent: "it worked" is a claim, and a claim is what you get when nobody wrote down what was read, in what order, to reach it.
Chain of Consciousness records that. Not a score an agent reports about itself, but the trace a third party can walk: the inputs, the sequence, the decision. It is the difference between a pilot you believe and a pilot you can grade.
pip install chain-of-consciousness and npm install chain-of-consciousness