← Back to blog

What's the Actual ROI of Deploying an AI Agent vs. a Human?

The slide says 15 to 300 times cheaper. The honest ledger has four lines, and it starts negative.

Published July 2026 · 9 min read · ROI / agent economics / reliability / oversight


In the fall of 2024, Salesforce put a price on a piece of digital labor: two dollars. Agentforce, unveiled at Dreamforce and launched that October, would handle a customer conversation end to end for $2 a pop, and Marc Benioff spent the following year describing the product line as a digital labor revolution. The pitch was irresistible arithmetic: a loaded support rep costs thousands a month; conversations cost two dollars; you can do the division yourself.

Within eight months, two things happened, both instructive. In May 2025, Salesforce introduced new flexible pricing for Agentforce, moving away from the pure per-conversation model, which had closed roughly 5,000 deals of which only about 3,000 were paid, and had, in SaaStr's phrase, priced out anyone without a large AI budget. And in June 2025, Salesforce's own AI research group published CRMArena-Pro, a benchmark of nine frontier models on realistic CRM work built on Salesforce's own schemas. The result: agents succeeded on 58 percent of single-turn tasks, and 35 percent of multi-turn ones.

Sit with that pairing. The company selling the two-dollar conversation measured, itself, that when the conversation takes more than one step, the agent completes it about a third of the time. Nobody lied. The price is real and the benchmark is real. The actual ROI of an agent versus a human lives in the space between those two numbers, and almost nobody prices that space. That's what this essay is for.

The slide that writes itself

Search for "AI agent cost" and you will find a genre: vendor cost-guides and calculators that converge on the same comparison. A fully loaded human employee runs somewhere around $5,000 to $15,000 a month; a capable AI agent subscription runs $10 to $500. The headline practically formats itself: agents are 15 to 300 times cheaper. I want to be straight about sourcing here, because this is a numbers essay and the provenance of the numbers is the story: those figures come from the people who sell agents. They are ranges from vendor guides, not audited economics. But the per-unit claim is not really the problem. Tokens genuinely are cheaper than salaries, by orders of magnitude. The problem is that the comparison prices the wrong two things.

Here is the puzzle the slide cannot explain. Deloitte's Tech Trends 2026 reports that only 11 percent of organizations have multiagent systems in production, against 38 percent experimenting. MIT's State of AI in Business 2025, which I have quoted before and will keep quoting because it remains the most bracing accounting of the era, found 95 percent of enterprise gen-AI initiatives producing zero measurable P&L return on $30 to 40 billion of spend. If agents were actually 15 to 300 times cheaper than the humans doing the same work, those numbers would be inexplicable. CFOs do not leave 300x savings on the table for governance reasons. The savings, as advertised, mostly are not there to take. Three omissions explain why.

The three costs the slide leaves out

First: most firms pay for both. The comparison assumes the agent replaces the human. Overwhelmingly, it does not; it runs alongside them. MIT's report has a verbatim line for this that deserves wide circulation: across sectors it found "significant pilot activity but little to no structural change," a pattern it summarizes as "widespread experimentation without transformation." Deloitte's 2026 agentic-AI guidance is written from the same observation, urging organizations to actually redesign roles and workflows around what it calls a silicon-based workforce, advice that only needs giving because it is rare. If the human stays at full scope, the agent's price is nearly irrelevant to net cost: you did not replace a salary, you added a subscription on top of one. The 15-300x slide compares the agent to the human it was supposed to remove. The honest ledger compares the agent plus the human you kept to the human alone, and that ledger starts negative.

Second: the oversight tax. Making an agent safe to deploy costs money, and the safer you need it, the more you spend. Vendor TCO guides themselves, read carefully, admit this: human-in-the-loop review workflows reportedly add 15 to 20 percent to build cost, and the hidden line items, integration, monitoring, maintenance, infrastructure, are commonly estimated to equal or exceed the subscription itself, with guidance to budget 50 to 100 percent on top. Again: vendor-published ranges, cited as such. But notice the shape rather than the digits. The oversight tax is not overhead in the ordinary sense; it scales with distrust. Every approval gate, every review dashboard, every audit trail is a purchase of reliability the agent does not natively have. Which means the cost advantage and the safety of the deployment are the same budget, spent against each other.

Third: the sticker is a quarter of the invoice. The same TCO literature puts initial development at 25 to 35 percent of three-year cost of ownership, meaning the $80,000 build in the proposal implies something like $230,000 to $320,000 over three years once inference, infrastructure, maintenance, and oversight are counted, and reports enterprises under-budgeting by 40 to 60 percent. Vendor numbers, illustrative only. But the direction matches everything else in this essay: the visible number prices the token; the real number prices the system that makes the token useful.

Why you can't fire the supervisor

The three omissions share one root, and it is the same root I keep hitting from different directions this year: reliability. An agent that executes a workflow step correctly 85 percent of the time completes a ten-step workflow about 20 percent of the time, because reliability compounds against you. I have called this the 85 percent problem, and its financial expression is simple: you cannot remove the human supervisor of an agent you cannot trust unattended. The supervision is not a transition cost that fades; at current reliability it is a permanent operating cost, and it is a human one, priced in salary.

Salesforce's own benchmark makes this concrete better than any vendor deck. CRMArena-Pro found agents reasonably strong, 83 percent, on structured, well-bounded workflows, then dropping to 35 percent on multi-turn tasks requiring context and follow-through, with a bonus finding that should chill anyone deploying near customer data: confidentiality awareness was poor, and prompting the agent to be more careful about it measurably reduced task accuracy. There is the oversight tax, derived experimentally by the vendor itself. The 65 percent of multi-turn tasks that fail do not fail silently and free; they fail into the lap of a human whose job is now cleanup and review, a job that did not exist before the agent arrived.

This is the money chapter of an argument I made in "the year of the demo": demos are short and production is long, so per-step unreliability that a demo survives, a workflow does not. The ROI version is: the demo shows you the token price, and production bills you for the reliability gap.

Both sets of numbers are true

Now the reconciliation, because the data genuinely splits in two, and the split is the answer.

The optimistic half is not fabricated. IDC reports organizations averaging a 2.3x return on agentic AI investments within about 13 months. PagerDuty's 2025 survey found more than half of companies had deployed AI agents in some form. And the narrow wins underneath such numbers are real: high-volume document processing, structured data extraction, first-pass triage, code assistance with tests as the verifier. The grim half is equally true: 11 percent in production, 95 percent of pilots without P&L impact, benchmarks failing two-thirds of multi-step work.

These halves coexist because they describe different tasks, not different truths. The ROI of "an agent versus a human" is not a number; it is a crossover curve, and the crossover is task type. An agent deployment clears the curve when four things line up: the task is high-volume, the judgment content is low, the output is cheap to verify, and, the one everyone skips, the human labor is genuinely removed or redeployed when the agent arrives. Miss the first three and you pay the reliability tax in cleanup. Miss the fourth and there is no return by construction, because nothing was ever subtracted from the cost side. The firms in the successful few percent are not smarter buyers of the same product; they are running different tasks through it, tasks on the winning side of the curve, and they redesigned the workflow so a subtraction actually occurred. Anyone who quotes you only the rosy half or only the grim half is, respectively, selling you something or excusing something.

The worksheet

Here is the practical residue, compact enough to use in the next budget meeting.

Price the deployment with four lines, not one. Net ROI = (the human cost actually removed or redeployed) minus (subscription and inference) minus (the oversight line: who reviews the agent's output, at what loaded rate, for what fraction of their time) minus (the TCO remainder: integration, monitoring, maintenance, which the vendor literature itself says can match the subscription). If line one is zero, stop: whatever you are buying, it is an augmentation, and it should be justified as one, on quality or speed, not sold internally as headcount economics.

Run it once on a real case and you will feel the slide dissolve. Take a five-person tier-1 support team at a mid-range loaded cost of $7,000 a month each, $35,000 of monthly labor, and an agent quoted at $2,000 a month all-in for subscriptions and inference. The slide says you just saved $33,000. Now the four lines. Suppose the agent takes 60 percent of ticket volume, but at something like Salesforce's own measured multi-turn reliability, a third of its multi-step cases come back for human rescue, so the team's real workload drops by roughly 40 percent, and, because you cannot staff 3.0 people for spiky coverage, you actually reduce or redeploy two: line one is $14,000. Line two is $2,000. Line three: a senior agent now spends half their time on review queues and escalation cleanup, call it $4,500 at their loaded rate. Line four, on the vendor literature's own guidance that hidden costs roughly match the subscription: another $2,000. Net: about $5,500 a month, a real and respectable return near 1.4x on spend, and nowhere within sight of 15 to 300 times. And that arithmetic assumed the reorganization actually happened. Leave the team at five, as the unredesigned majority do, and the same deployment nets negative $8,500 a month, indefinitely, while every engagement dashboard shows the agent busily succeeding. Both spreadsheets contain a working agent. Only one contains a return, and the difference is nothing the model did.

One caution on "redeployed," because it is the word where ROI goes to hide: redeployed labor only counts as a return if the new work is something you would have hired for anyway. Moving two support reps to "quality initiatives" nobody budgeted is not redeployment; it is a raise for the slide's author.

Before deploying, ask the crossover questions in order. Is the volume high enough to amortize the setup? Is the judgment content low enough that 85 to 95 percent per-step reliability doesn't compound into cleanup? Is verification cheap, a test suite, a schema check, a total that must reconcile, so that catching the failures does not itself cost a salary? And will a human's scope actually shrink on a schedule someone will sign? Four yeses is a deployment. Three is a pilot. Fewer is a demo.

And give the pilot a P&L line from day one. MIT's 95 percent figure is not a measure of failed technology; it is a measure of unmeasured deployments, initiatives that tracked engagement and saved-minutes and never once computed removed cost. Saved minutes that do not become removed or redeployed cost are an anecdote. The two-dollar conversation was never the number that mattered. The number that matters is what it costs to trust the conversation, and until the reliability curves move, that number is denominated not in tokens but in the salary of the person still reading over the agent's shoulder. Price the shoulder. That's the actual ROI math, and once you run it, both the 11 percent and the enthusiasm of the winners stop being mysterious: they are the same curve, read from opposite sides.


Sources

The oversight tax is the cost of trusting the agent's output. Make that cost cheaper and the ledger changes.

Line three of the worksheet, the human still reading over the agent's shoulder, is really the cost of verifying what the agent did. The cheaper and more mechanical that check gets, the more deployments clear the crossover curve. The agent trust stack is that layer built as installable pieces: provenance of what an agent actually did, ratings of how much its output is worth, and verification that the constraints held, so the shoulder you have to price is a schema check instead of a salary.

Read the Theory of Agent Trust

pip install agent-trust-stack  ·  npm install agent-trust-stack

Or the pieces on their own: pip install chain-of-consciousness / npm install chain-of-consciousness for the provenance record of what an agent did.