In 1969, a Greek banker named Minos Zombanakis had a problem shaped like $80 million. His bank, Manufacturers Hanover, was arranging a syndicated loan to the Shah of Iran, and a loan that size, split across a group of banks in a floating-rate world, needed a reference rate everyone could live with. His solution was pragmatic to the point of improvisation: price the loan off what a handful of reference banks reported their own funding costs to be. Ask the banks. Write down the answers. Average them.
That convention got a name, and in 1986 the British Bankers' Association formalized it for the dollar, sterling, and yen: the London Interbank Offered Rate. By 2021, secondary tallies put more than $230 trillion of financial assets referencing LIBOR: mortgages, student loans, corporate credit, and a derivatives market that dwarfed all of it.
Here is the part to sit with: nobody ever decided LIBOR should carry that weight. There was no summit, no standards body vote, no moment where someone said "this survey is now the load-bearing number of the global financial system." It accreted, one contract at a time, because it already existed and existing is the strongest qualification a reference can have. Every engineer who has watched a quick script become tier-one infrastructure knows this shape. LIBOR is what happens when the dependency nobody chose becomes the dependency nobody can remove, at planetary scale, denominated in trillions.
Each day, panel banks were asked: at what rate could you borrow funds, were you to do so?
Not: what did you pay this morning. Not: show us the trade. What do you estimate you could pay, if you were to ask. No transaction was required to produce a submission. The world's most referenced number was, mechanically, a daily opinion poll of interested parties about their own creditworthiness.
Before you laugh at 1969, concede what the designers were solving, because the concession makes the story better, not worse. Borrowers want a forward-looking term rate: a number that tells you today what your payment will be for the next three months. You cannot read a three-month forward rate off yesterday's overnight transactions; the trades that would define it barely exist as a continuous market. The survey was not laziness. It was the only technology anyone had for manufacturing a forward-looking number, and for decades, in a clubby interbank market where the banks genuinely did lend to each other at quotable rates, it worked well enough that nobody looked underneath.
The estimate wasn't a shortcut around available data. It was a bridge over data that didn't exist. Hold that thought, because the same defense is being offered for other numbers right now.
The scandal that broke in 2012 (Barclays settled first, for $450 million; UBS later paid $1.5 billion and RBS $612 million, part of a wave that ultimately reached billions across more than a dozen institutions) is usually told as one story. It was two, running in opposite directions, and the less famous one is the more instructive.
Direction one, upward, for money: traders asked submitters to nudge the day's number a basis point or two to favor derivative positions. Greedy and simple. This is the half everyone remembers.
Direction two, downward, for reputation. Between 2007 and 2009, with the financial crisis metastasizing, Barclays submitted what regulators later called dishonestly low rates — not to profit on a trade, but to dampen speculation about its own viability. A high LIBOR submission says other banks charge me more because they trust me less. During a solvency panic, submitting your honest number is indistinguishable from announcing your own distress. So banks lowballed, to look healthy.
Notice what the second direction actually is. It is not theft; no trader pocketed the spread. It is a feedback loop: the signal was corrupted by the thing it was supposed to measure. The survey asked banks to publish a daily self-assessment whose honest value could trigger the very run it reported the risk of. A metric that punishes honest reporting will not get honest reports. That is not a statement about bankers; it is a statement about metrics. Any engineer who has watched teams game a dashboard the moment it started driving promotions has seen the same loop with smaller numbers.
For a decade, the human face of LIBOR was Tom Hayes, the UBS and Citigroup trader jailed in 2015. He was the scandal's proof of intent, the reason the whole affair could be filed under "crooked bankers."
In July 2025, the UK Supreme Court quashed the convictions of Hayes and Carlo Palombo. The ruling was about the trials, not a declaration of innocence: the judges' directions had, in the court's view, effectively prevented juries from considering the key question of whether the traders acted dishonestly. Be precise about what that means, because both overstatements are wrong. It does not undo the banks' settled findings of manipulation; the institutional record stands. But the criminal question (did these individuals act dishonestly, as a jury sees it) was, per the Supreme Court, never properly put.
For this essay's purposes, the quashings are clarifying. Strip the villains out entirely, and what remains? A number that moved when banks feared for their reputations. A number a submitter could shade with no trade to contradict him. A number whose honest reporting was punished by the market it fed. The number was broken whether or not anyone was dishonest. The design did not require crooks to fail; it merely made failure unattributable when it came. If your system's integrity story depends on prosecuting individuals afterward, you do not have an integrity story. You have a deterrence story, and in this case even that half-collapsed on appeal, thirteen years later.
The repair took a decade. The scandal broke in 2012; the final USD LIBOR panel settings ceased on June 30, 2023, after which no bank was required to submit. Eleven years, for a number everyone influential had agreed was rotten, with regulators on multiple continents pushing in the same direction. Anyone expecting a broken benchmark to be fixed in a product cycle should study that timeline.
What replaced it, for dollars, was SOFR: the Secured Overnight Financing Rate, computed by the New York Fed from actual overnight repo transactions collateralized by US Treasuries. The comparison that matters fits in two lines:
LIBOR: transactions required to produce a submission: zero. SOFR: average daily transaction volume underlying the rate since April 2018: roughly $977 billion, regularly above a trillion dollars a day.
The replacement for an estimate was not a better estimate, a bigger panel, or an audited survey. It was a trillion dollars a day of observed trades. That is the whole migration, made numeric: from "ask the participants" to "measure the market," from a number someone produces to a number someone computes.
Two honest footnotes, so this doesn't read as a clean-success story. First, SOFR did not fix LIBOR; it replaced it with a different instrument that answers a narrower question well instead of a broader question badly. SOFR is secured and overnight; LIBOR was unsecured and forward-looking, and the forward-looking term structure borrowers loved had to be rebuilt on top with derived term rates, a real and continuing complaint from the loan market. The migration traded convenience for verifiability, on purpose. Second, this is the USD story; sterling, euro, and yen each got their own transaction-based successors (SONIA, €STR, TONA) on the same principle.
Now the present-day read, and a correction to the analogy people usually reach for.
When the AI field worries about its own capability numbers (and it does worry; there is a small literature on it now, with papers titled things like "Benchmarking is Broken: Don't Let AI be its Own Judge") the financial comparison that gets drawn is to Moody's and the credit-rating agencies. It is the wrong precedent. Credit ratings fail through an issuer-pays conflict: the rated party hires the rater, and the disease is in who signs the check. The benchmark problem is different. A lab that runs its own evals and publishes its own scores has no rater at all. The number is a self-report, produced by the interested party, with no underlying transaction anyone else can inspect, no observable event that forces the number to be what it is. That is not Moody's. That is a LIBOR submission: at what score could your model perform, were it to be tested?
And the incentive structure runs both LIBOR directions, faithfully. Upward for money: capability claims sell APIs and raise rounds, the trader's nudge. And the subtler one: a lab's published numbers are also read as solvency signals in a talent and capital race, which is the Barclays crisis dynamic: the reported number shading toward what the reporter needs the world to believe about its health.
Here is the encouraging part, and the reason the history matters rather than merely rhyming. The benchmark-reform literature has independently converged on LIBOR's own cure. One proposal recasts evaluation as "a standardized, proctored examination rather than an 'open-book' contest of self-reported scores." Read that against the SOFR transition and it is the same move, reinvented from scratch by people who weren't writing about finance: replace the self-report with an observed process: a transaction, or a proctored run, or any event a third party can independently recompute. The playbook exists. It has been executed once, at $230 trillion scale, against incumbents with every incentive to stall. The second field running it doesn't have to rediscover each step.
Portable beyond finance, and beyond AI.
Audit your load-bearing numbers for the LIBOR property. For each metric your organization treats as ground truth (uptime, conversion, model accuracy, velocity) ask one question: does a transaction discipline this number, or does an interested party produce it? Self-reported numbers are not lies. They are unpriced risk, and the risk compounds with every contract, decision, and bonus you index to them. The $230 trillion did not accumulate because anyone trusted the survey deeply. It accumulated because nobody re-asked the question after the number became convenient.
Watch for the feedback direction, not just the greed direction. The manipulation everyone anticipates is upward, for gain. The one that corrupted LIBOR at the systemic moment was downward, for survival: honest reporting punished by the market reading it. Any metric that doubles as a health signal for its own reporter (error budgets before a launch review, incident counts that feed a team's standing, capability scores mid-fundraise) carries this loop by construction. Design for it before the crisis, because during one, every reporter discovers the same arithmetic Barclays did.
And budget a decade. The migration from estimate to transaction is possible — that is the genuinely hopeful lesson of 2012–2023 — but it took eleven years with global regulators aligned and the old number publicly disgraced. If you are waiting for capability benchmarks, or any entrenched self-report, to be replaced by observed process, the precedent says: it happens, it has a playbook, and it is slower than the people demanding it can stand. Start yours before your 2012, not after.
Zombanakis, by every account, was solving Tuesday's problem: one loan, one Shah, one rate the syndicate would accept. The survey was a good answer to that question. It was never an answer to the question fifty years of accretion eventually asked of it. The most useful thing you can do with his story is go find the number in your own systems that is quietly being asked a question it was never designed to answer, while it still costs one meeting to fix instead of eleven years.
Sourcing notes: the widely cited "roughly $9 billion in industry fines" aggregate is deliberately not printed — the three individually sourced settlements here (Barclays $450M, UBS $1.5B, RBS $612M) total about $2.6B, and the larger figure lacked a named tally at research time, so the essay says "billions across more than a dozen institutions." The $230 trillion exposure figure is from secondary coverage keyed to 2021 and moves with the year quoted. The 2025 quashings are written per the research bound: a ruling on trial fairness, not a finding of innocence, leaving the banks' settled findings undisturbed. The Zombanakis origin details (the $80M loan, Manufacturers Hanover, the Shah) are consistent across secondary accounts; no primary was reached. SOFR's ~$977B average daily underlying volume since April 2018 is as reported by the transition's official materials; the term-rate complaint against SOFR is acknowledged in-text rather than resolved.
The LIBOR-to-SOFR migration is one sentence long: stop asking an interested party to produce the number, and start computing it from events a third party can see. The question to carry into any system is the same one. Is this figure produced, or is it computed from a record someone else can recount?
Chain of Consciousness keeps that record for agent decisions: what was read, in what order, to arrive at what was claimed. It is the difference between a submission and a transaction.
pip install chain-of-consciousness and npm install chain-of-consciousness