Uber Burned Its 2026 AI Budget in Four Months, and the $1,500 Cap Is Per Tool

In the spring of 2026, Uber's chief technology officer sat down for a personal demo session with Claude Code, the agentic coding tool the company had rolled out to its engineers the previous December. The session ran two hours. It cost $1,200 in tokens.

That figure comes from The Information's April reporting, which is paywalled; it reaches this essay through Forbes' account of it, and so does every adoption and cost number below. Everything here dates from April to June 2026. Treat the chain as stated. What makes the $1,200 worth opening on is not the size. It is that the number exists at all, at that precision, for a two-hour block of one executive's afternoon. It is the kind of number Uber turns out to be excellent at producing: a cost, measured to the dollar, attached to nothing.

An undisclosed budget, exhausted at an undisclosed rate

The same reporting carried the claim that made the rounds: according to CTO Praveen Neppalli Naga, Uber's AI budget for all of 2026 was gone roughly four months into the year.

Almost every write-up missed the most important fact about that sentence: nobody outside Uber knows what the budget was. Uber has never disclosed it. Aggregators filled the blank with the biggest number in the company's filings, $3.4 billion, and headlined a token crisis of that size. But $3.4 billion is Uber's entire research-and-development expense for fiscal 2025, straight off the SEC-filed income statement: salaries, infrastructure, every project the company builds, a line that grew about nine percent over the prior year for reasons that long predate any coding agent. Pinning the token story to it relabels the whole engineering organization as an inference bill. The honest version is less quotable. An undisclosed budget was exhausted at a partially disclosed rate, and the only people who know the denominator are not saying.

Hold that word, denominator. It is the whole essay.

The COO asks what the tokens bought

On May 22, Uber president and COO Andrew Macdonald sat for the Rapid Response podcast and got the question every finance organization is now asking. His answer, as Fortune transcribed it: "That link is not there yet." The link he means is the one between token spend and shipped product. The fuller quote, same transcription, brackets Fortune's: "If you're not actually able to draw a direct line to how [many] useful features and functionality you're shipping to your users, that trade becomes harder to justify."

A COO asking what the money bought is not news. What earns this one an essay is the company it happened at. Uber is not some under-instrumented startup discovering observability. Its cost telemetry on AI coding is, on the public record, unusually good. And every instrument points the same way.

What Uber can measure, and what it cannot

Here is the measurement inventory, all of it from the same April reporting chain. Claude Code adoption went from 32 percent of engineers in February to 84 percent classified as agentic coding users in March. Roughly 95 percent of engineers were using AI tools monthly by spring. Roughly 70 percent of committed code originated from AI tools. Eleven percent of live backend updates were written by agents with no human oversight. Monthly cost per engineer ran $150 to $250 on average, $500 to $2,000 for power users. And the CTO's own demo: $1,200, two hours.

Read the list again and sort it. Dollars per engineer, dollars per session: cost. Adoption share, monthly-use share: volume. Share of committed code, share of backend updates: volume again, of a particular kind we will get to. There is not one number in the inventory for what came out the other side. Features shipped against the pre-agent baseline. Cycle time on comparable work. Incidents, revenue, anything. Uber instrumented the numerator of its AI investment to the dollar and never instrumented a denominator, and a company cannot discover a ratio it never built a denominator for. So when Macdonald says the link is not there yet, the precise reading is not that Uber lacks data. It is that no amount of reading these particular instruments will ever produce the answer he is asking for, because none of them is pointed at the thing he is asking about.

This is the opposite failure from the one the named-AI-thing genre usually covers. When we wrote about Devin, the "first AI software engineer" that failed 86 percent of its benchmark tasks, the story was a tool that underdelivered on capability. Uber's tools overdelivered on demand: engineers liked them enough to triple adoption in a month and blow through a year of budget in four. Useful and unmeasured turns out to be its own failure mode, and a more expensive one per week than useless ever was.

The 70 percent is not what it is quoted as

The number that circulates as if it settled the value question is the 70 percent of committed code originating from AI. It gets quoted as proof the spend is working. Look at what it actually measures: of the code that reached a commit, what share a model produced rather than a human. That is an input substitution rate. It tells you how much of the typing got outsourced. It does not tell you whether more shipped than last quarter, whether anything shipped sooner, or whether a user somewhere got a feature they would not otherwise have gotten. A company where AI writes 70 percent of committed code and ships the same roadmap on the same dates has automated its typing, not its delivery.

That is the category error in one line: an input metric wearing an output metric's clothes. Once you see it there, you see the whole inventory the same way. Every impressive number in the Uber story describes what went in.

The fix repeats the mistake

On June 2, Bloomberg's Natalie Lung reported Uber's response, from a company spokesperson. The passage, quoted in full by Simon Willison, deserves a close read:

"The rideshare giant is limiting all employees to $1,500 in monthly token spending per AI coding tool, an Uber spokesperson said in response to a Bloomberg News inquiry. That means spending on one tool doesn't have a bearing on the budget for another. The limits, which have been instituted in recent months, only apply to agentic coding software such as Cursor or Anthropic PBC's Claude Code."

The headline reading is that Uber capped AI spending at $1,500 a month. The text says something different. The cap is per tool, and spending on one tool has no bearing on the budget for another. A per-tool cap is not a spend cap; it is a spend cap multiplied by the vendor count. With two named agentic tools in the sentence, the actual ceiling is $3,000. Approve a third tool and the ceiling rises to $4,500 without anyone making a decision called "raising the ceiling." The control loosens through procurement, which is the one function guaranteed to keep adding tools.

Set the loophole aside and the deeper problem is unchanged. The cap rations the input more precisely than before. It measures the output exactly as much as before, which is not at all. Uber's answer to "we cannot connect spend to value" was to control spend harder, which is the available lever, and I understand why a company reaches for it. But it is the same ledger with a tighter clamp on the same side. Nothing about a cap, at any dollar level, per tool or global, produces the number Macdonald asked for in May.

Your ledger, probably

If you can quote your agent spend to the dollar and your agent value only as a share of code written, you have Uber's ledger at Uber's stage, whatever your scale. The cost side has its own well-documented omissions, and we have written that ledger before, in "What's the Actual ROI of Deploying an AI Agent vs. a Human?"; this is the other side of it. The value column is not missing because measuring it is impossible. It is missing because nobody made it a precondition, and adoption filled the vacuum as adoption does.

The first honest move is not a cap. It is picking one output number before you touch the spending dial: features shipped per quarter against your pre-agent baseline, lead time on comparable tickets, escaped defects per release, one number your business already believes in. Instrument that, let it run long enough to mean something, and then a cap becomes a policy you can defend with a ratio instead of a feeling of control. If the ratio never materializes, that is also an answer, and a cheaper one to learn on purpose than by April.

Uber can tell you, to the dollar, what two hours of Claude Code costs. What two hours of Claude Code makes, nobody at Uber can say. Their own president said so, into a microphone: that link is not there yet.


Sources