← Back to blog

"Best Execution" Is the 90-Year-Old Legal Template for Agent SLAs

The name sounds like a promise that your broker gets you the best price. It is not. It is a promise about process, not outcome, and that distinction is the single most useful template for writing an enforceable contract for a probabilistic AI agent.

Published July 2026 · 12 min read · agent SLAs / accountability / securities law / calibration


"Best execution" is one of the oldest obligations in securities law, and its name is a lie of the most instructive kind. It sounds like a promise that your broker will get you the best available price. It is not. Under FINRA Rule 5310, a broker can route your order somewhere that fills it at a worse price than the best quote showing on another exchange, on that specific trade, and remain in perfect compliance. What the rule actually demands is not a good price but a good process. In its own words, a broker must "use reasonable diligence to ascertain the best market for the subject security and buy or sell in such market so that the resultant price to the customer is as favorable as possible under prevailing market conditions." Reasonable diligence. Not a guaranteed outcome.

That distinction, between promising an outcome and promising a process, is about ninety years old, and it is the single most useful template in existence for anyone trying to write an enforceable contract for a probabilistic AI agent. The whole anxious industry conversation about "AI SLAs," about what you can possibly promise for a system that is wrong some unpredictable fraction of the time, has a worked precedent sitting in plain sight, in the rulebook that governs how your stock trades get filled. And lest anyone think a process standard is a soft one: in December 2020 the U.S. Securities and Exchange Commission fined the brokerage Robinhood sixty-five million dollars, in part for failing exactly this duty. A process standard, done right, has teeth.

The problem the rule was built to solve

Start with why the law ever settled for a process instead of an outcome, because it is the same reason you cannot write "correct every time" into a contract for a language model.

A broker does not control the price. The market does. Prices move between the instant an order is placed and the instant it is filled; liquidity appears and vanishes; the best-looking quote can be stale or unreachable. A broker operating in good faith, doing everything right, will still get a worse-than-theoretical-best price on plenty of individual trades, through no fault of its own, because the system it operates in is fundamentally probabilistic and partly outside its control. If the law had demanded a guaranteed best price on every trade, it would have demanded the impossible, and a rule that demands the impossible is either ignored or gamed into meaninglessness.

So securities regulation never demanded it. It demanded diligence, and it demanded review. And look at how precisely that maps onto a probabilistic agent. Your LLM does not control whether it retrieves the right document, whether the tool it calls returns clean data, whether the question it was handed is even answerable. It operates in a system that is fundamentally probabilistic and partly outside its control, and it will be wrong on some fraction of individual calls no matter how well it is built. Promise per-call correctness and you have promised the impossible; the SLA becomes a lie or a metric so hedged it means nothing. This is not a new problem. It is the broker's problem, and the broker's problem was solved a long time ago.

What the rule actually requires

The solution has two halves, and both transfer directly.

The first half is the diligence standard, and Rule 5310 makes it concrete by naming the factors a broker's process has to weigh: the character of the market for the security, including its price, volatility, and liquidity; the size and type of the transaction; the number of markets checked; the accessibility of the quotation; and the terms and conditions of the order. Notice what that is. It is not a formula that outputs the right venue. It is a list of considerations a defensible decision procedure has to take into account. Compliance is judged on whether your process seriously weighed the right things, not on whether it happened to land on the best price this time.

The second half is the part that does the real work, and it is explicitly a population-level obligation. A firm that routes its orders elsewhere, the rule says, "must have procedures in place to ensure the member periodically conducts regular and rigorous reviews of the quality of the executions" it is getting. In practice that means, at least quarterly, on a security-by-security and order-type basis, comparing the execution quality it gets through its current routing against what competing venues would have delivered, examining price improvement and disimprovement, speed, likelihood of execution, and cost, documenting the analysis, and modifying the routing when the review turns up material differences. And the rule frames this whole apparatus with a single clause that is worth reading twice: a firm must conduct this regular and rigorous review "if it does not conduct an order-by-order review."

Sit with that. The regulation explicitly contemplates that you will not check every order against the theoretical best. It offers, as the alternative, a systematic audit of the population. That is not a loophole in the rule. It is the design of the rule. The accountable unit was never the individual trade. It was the routing policy and the periodic review of how that policy performs across all the trades.

Why "ninety years"

A quick word on the vintage, because precision matters here and the round number hides some history. The specific rule I have been quoting, FINRA 5310, is modern; the first explicit best-execution rule came from the NASD, FINRA's predecessor, in 1968. But the duty is much older than any rule that codified it. It descends from a broker's common-law obligation as an agent to act in the customer's interest, a duty courts were already articulating in the nineteenth century, and it became a fixture of federal securities regulation with the framework built on the Securities Exchange Act of 1934, enforced through that Act's antifraud provisions. Call it roughly ninety years old as a load-bearing piece of federal securities law, and older still as an idea. The point is not the birthday of any particular rule. It is that the market confronted the problem of holding someone accountable for good conduct inside a probabilistic system nobody fully controls, and it worked out a durable answer almost a century before anyone needed it for software.

Calibration is order routing

Here is the mapping, and it is exact rather than loose. Everything a well-built agent does when it is uncertain, the whole calibration toolkit of route, retry, verify, abstain, escalate, is order routing. "Send this order to the venue most likely to fill it well" and "send this query to a stronger model, or to a tool, or to a human, or refuse to answer" are the same kind of decision: a policy choosing, per the situation, how to handle the thing in front of it.

So the two halves of best execution reassemble, feature for feature, as the two halves of an honest agent contract. The broker's five diligence factors become the agent's: its confidence in the answer, the stakes of being wrong, the cost and latency budget, the availability of a tool or a human to check with. A defensible agent policy weighs those the way a defensible routing policy weighs price and liquidity and accessibility. And the regular-and-rigorous review becomes an audit of the agent's decisions over a population of calls: its accuracy, yes, but also its abstention rate, the quality of its escalations, and its calibration, whether the confidence it reports actually tracks how often it is right. Broken out by task type and stakes. Compared against the alternatives it could have routed to. Documented. And, crucially, used to modify the policy when the numbers move. That is the whole shape of an agent-quality program, and it is a ninety-year-old compliance rule with the nouns swapped.

The measurement side of this, the case that you should evaluate agents on populations rather than by connoisseurship over individual outputs, we have written about separately, and it is the necessary companion. But measuring on populations tells you how to know whether a process is good. Best execution tells you something the measurement argument does not: that a good-process-audited-over-a-population standard is a legitimate, enforceable form of accountability, recognized as such by law, and not a way of wriggling out of one.

It has teeth, and it has a known weakness

Which brings us back to Robinhood, because the objection to everything above is obvious and needs answering. A process-not-outcome standard can sound like an elaborate excuse: if nobody has to be right on any particular call, accountability seems to have quietly evaporated.

It has not, and the securities markets are the proof, because the process standard is enforced, hard, at the level of the policy and the review. The SEC's 2020 order against Robinhood found that the firm had failed to satisfy its duty to seek the best reasonably available terms for its customers' orders, and that it had misled those customers about how it made its money: through payment for order flow, the practice of routing orders to the wholesale trading firms that pay for the privilege. The regulator did not build its case by parading a list of individually botched trades. It attacked the routing process and the misrepresentation of it, and it quantified the harm at the population level: the inferior prices, the order found, cost customers about $34.1 million. The penalty was sixty-five million dollars, and the remedy included bringing in an independent consultant to review the firm's order-routing policies and procedures. That is accountability landing exactly where an agent vendor's should land, on the policy and the review, and it is not gentle.

But the Robinhood case is also the honest edge of this whole template, because payment for order flow is precisely where best execution is most contested. When the venue that pays you to route to it is also the venue you are supposed to evaluate as "best," your reasonable diligence is fighting a conflict of interest, and the entire question of whether "regular and rigorous review" can genuinely police that conflict, or whether it becomes a rubber stamp for the routing you were paid to prefer, is a live and unresolved argument, the one that erupted into public view during the 2021 meme-stock episode. The template is not perfect. It has a specific failure mode: the review can be captured by the very incentive it exists to check.

And that failure mode transfers to agents with zero translation. An internal "quality review" of your agent, run by the team whose cheaper, faster model the review might indict, is Robinhood's conflict wearing a company lanyard. If the people auditing whether the agent should have escalated are the same people rewarded for the agent not escalating, you have built the process standard's known weakness directly into your accountability. Best execution does not just hand you the template; it hands you, in the same package, the exact way the template gets gamed.

The contract you can actually write

So here is the practical residue, the agent SLA drawn straight off Rule 5310.

Stop trying to promise per-call outcomes. A blanket "99.9% correct" for a probabilistic agent is either a falsehood or a number so gently defined that it certifies nothing. Best execution never promised the best price, and it has governed trillions of dollars of trades for decades. Do not promise the best answer.

Promise the policy. Specify, in the contract, the decision procedure: how the agent routes, retries, verifies, abstains, and escalates, and the factors that drive those choices, its confidence, the stakes, the cost and latency budget, the tools and humans available to it. A documented, defensible policy is a real obligation, the way reasonable diligence is a real obligation, and a conflicted or thoughtless policy is a breach even if a given answer came out fine.

Promise the review. Commit to a regular and rigorous audit of the agent's behavior over a population of calls, on a schedule, broken out by task type and stakes, compared against the alternatives, measuring accuracy and abstention and escalation quality and calibration, documented, and with a standing duty to change the policy when the review shows drift. This is the enforceable heart of the thing. It is what makes the whole arrangement accountability rather than hand-waving.

And protect the review from capture, because that is the one lesson the securities markets learned the hard way and you can learn for free. The party auditing the agent must not be the party whose corners the audit would expose. Independent review is not a compliance nicety; it is the component that gives the entire standard its teeth. Without it, "regular and rigorous review" becomes a quarterly ceremony that always concludes the cheap model was fine.

The deepest thing best execution has to teach is a reframe. "We cannot guarantee every answer" sounds, to a nervous buyer or a nervous vendor, like a confession that no one is accountable. It is the opposite. It is the honest first sentence of the only accountability model that has ever actually worked for a probabilistic system that no one controls. The market built that model ninety years ago for a machine far larger and stranger than any language model, the market itself, and the answer it reached was never to guarantee the trade. It was to govern the policy and audit the population, and to punish, in real money, the firms whose policies were bad or whose reviews were captured. Your agent needs precisely that, and nothing more exotic. Write the policy. Audit the population. Keep the reviewer honest. That is what accountability for a probabilistic system looks like, and it has looked like that since before there were computers to need it.


Sources

Write the policy. Audit the population. Keep the reviewer honest. That is what an enforceable contract for a probabilistic agent actually looks like.

Best execution hands you the shape of agent accountability, and the agent trust stack is the machinery for building it: a documented, tamper-evident record of what the agent's policy actually did on each call (the provenance layer), an audit of the agent's behavior over a whole population of calls scored against reality rather than against its own confidence (the ratings layer), and a check of the output against ground truth where the stakes demand it (the verification layer). The one thing the securities markets learned in real money is the part you cannot skip: keep the review independent of the party it would indict.

Read the Theory of Agent Trust

pip install agent-trust-stack  ·  npm install agent-trust-stack

Or the population-audit layer on its own: pip install agent-rating-protocol / npm install agent-rating-protocol.