The improvement you measured is a covariance, not a cause
In 2013, MD Anderson Cancer Center began building a clinical decision support tool with IBM, using Watson technology. It was called the Oncology Expert Advisor. It was supposed to recommend treatments, match patients to trials, and help evaluate cases. Over roughly four years the project consumed at least $62 million.
By September 2016 it was not in clinical use. IBM ended support for the pilot and demo systems effective September 1 of that year. A University of Texas System audit followed, and the press coverage arrived the next February.
The number everyone quotes is the $62 million. The fact that actually matters sits one line below it in the audit: the system had never been piloted anywhere outside MD Anderson.
That is not a statement about whether the software was any good. It is a statement about what could possibly have been known. A clinical decision support system trained on one institution's practice, evaluated only inside that institution, is being measured against the people who taught it. Agreement is the expected result. It would have been the expected result whether the recommendations were medically excellent or medically wrong, because the thing being measured and the thing doing the measuring share a source.
There was no out-of-sample test to fail. That is what the $62 million bought.
The audit is easy to misread, and misreading it is the fastest way to get this story wrong.
What the auditors found was a procurement problem. Contracts, approvals, compliance. Nearly all of the spending went ahead without board approval. They also noted that the tool could not exchange data with the Epic electronic health record system the hospital had moved to, which left oncologists consulting protocol and trial data from the system Epic replaced.
What the auditors explicitly did not do was evaluate the model. The review's own scope statement limits it to contracting and compliance and says it did not cover project management or system development. The 48-page document is not a verdict on whether the Oncology Expert Advisor gave good advice. Nobody produced that verdict, because producing it would have required the external pilot that never happened.
So there are two failures here and they are not the same failure. One is that money moved without the approvals it needed. The other is that four years of work generated no evidence that could travel outside the building. The first is a governance problem. The second is an epistemics problem, and it is the one that repeats everywhere.
Watson for Oncology, a related but distinct product, was trained with Memorial Sloan Kettering and did get deployed to other institutions. So we can see what happens when the test finally runs.
Concordance rates, meaning how often the system's recommendation matched the local tumor board, vary enormously by site and by cancer type. A double-blind study of 638 breast cancer patients in India reported 93% concordance for recommended or for-consideration treatment. Work presented at ASCO in 2017 covering 525 patients in Korea reported 73% for colon cancer and 49% for gastric cancer. A Korean study at Gil Medical Center put absolute concordance for colon cancer at 48.9%, rising to 65.8% when cases judged acceptable were included. For gastric cancer the same kind of analysis reported 41.5% at the recommended level and 87.7% at the for-consideration level.
Read those numbers as a group rather than individually. The same system, asked the same kind of question, agrees with local practice more than nine times in ten in one setting and fewer than half the time in another.
The interesting reading is not that the system was bad. It is that concordance was never measuring medical correctness in the first place. It was measuring the distance between one institution's encoded practice and another's. High agreement where local practice resembles the training institution, low agreement where it does not. That is a covariance between two things that share a cause, and it produces a number that looks like an accuracy score and behaves like a similarity score.
An organization that only ever measured this metric at home would see a high number every time and would learn nothing at all.
It would be comforting to file this under carelessness, but the internal number is persuasive for reasons that survive being smart and careful.
The first is that it is not imprecise. You can measure concordance at your own institution to as many decimal places as you like, and every one of them will be accurate. The precision is real. It is attached to the wrong quantity. High precision on the wrong estimand feels exactly like rigor from the inside, and it produces the confident charts.
The second is that the people doing the evaluating are the people who supplied the training signal. When the system disagrees with them, the natural reading in the room is that the system made a mistake, because in that room it did. Disagreement gets logged as error and corrected away. The evaluation loop is therefore not neutral about which direction the system moves. It rewards convergence on local practice, which is the same operation as destroying the system's ability to tell you anything you did not already believe.
The third is that nobody in the building has an incentive to produce the number that would embarrass everyone in it. An external pilot is the only party to the process that has no stake in the result. That is not a cynical point about human character. It is a structural point about where disconfirming evidence comes from, and it is why "we validated it internally" is a description of a procedure rather than of evidence.
If oncology feels far away, here is the version with a coefficient attached.
In 2024 a team at Scale AI built GSM1k, a fresh benchmark designed to mirror the style and difficulty of GSM8k, the widely used grade school math benchmark. The point was to ask a question that GSM8k can no longer answer about itself: when a model scores well, how much of that is arithmetic reasoning and how much is having seen the test?
Their paper, "A Careful Examination of Large Language Model Performance on Grade School Arithmetic," reports accuracy drops of up to 8% when models move from GSM8k to the fresh set. Several model families show systematic overfitting across nearly all sizes. And the finding that makes it more than a suspicion: a positive relationship, Spearman's r² = 0.36, between how likely a model is to generate GSM8k examples and how much its score falls on the held-out set. Models that can recite the test do worse when the test changes.
That coefficient is what turns an argument about principle into a measurement. A benchmark score is a joint measurement of capability and exposure, and nothing inside the score tells you the ratio.
While assembling this piece we ran into the thesis three separate times, in our own sources.
The GSM1k accuracy drop is reported in a good deal of secondary coverage as 13%. The paper's abstract says up to 8%. We took the primary.
The University of Texas audit is described in several retellings as posted on January 31, 2016, which cannot be right, because the events it audits run through September 2016. The posting date is January 31, 2017. A year fell off in transit.
The split of the $62 million between IBM and the consulting firm that supported the project is given slightly differently across accounts, roughly $39 to $40 million to IBM and roughly $21 to $23 million to the consultancy, while the total holds steady. This is why the figure here is "at least $62 million," which is the reporting's own hedge, rather than a precise sum.
None of these drifts is scandalous. Each write-up was doing its honest best to report what a study or an audit found. That is the point. A number does not need anyone to lie about it in order to arrive somewhere false. It only needs to be copied a few times by people who did not go back to the source, and each copy is individually reasonable.
The GSM1k paper also found that frontier models showed minimal signs of overfitting. They generalized to problems guaranteed to be absent from their training data.
That matters, and leaving it out would be the same sin this essay is about. The honest claim is not that benchmarks are theater or that measured progress is fake. Real capability gains exist and the same study that quantifies contamination also demonstrates them.
The claim is narrower and more useful: a measured improvement licenses a causal reading only when the thing you changed is the only thing that moved. When your evaluation shares a source with your training, whether that source is an institution, a market regime, or a scraped corpus, some fraction of the agreement you observe is an echo, and the score alone will not tell you how much.
Here is the version of this that nearly got past us this morning.
Working a research question on July 27, 2026, we wanted a data-grounded read on whether a particular kind of skilled work had been touched by AI adoption. The Anthropic Economic Index publishes a labor market impacts release with task penetration and job exposure data, so we pulled the task-level file and looked at every calibration and metrology task in it.
Every one read 0.0.
That is a striking result. Measured AI penetration into this work, zero. It would have made an excellent line, and we were most of the way to writing it.
Then we counted the rest of the file. By our count of that file, of 17,998 task statements only 1,354 carry a nonzero value. About 92.5% of the file reads 0.0. Our finding was the modal value of the dataset. We had measured the sparsity of a file and were about to report it as a fact about an occupation.
The fix was to move to the occupation-level file, where the distribution actually discriminates: by our count, a mean of 0.0770 across 756 occupations with 54.4% sitting at exactly zero. There the relevant proxies read 0.0324 and 0.0. Still low, but now meaningfully low, because there was a spread to be low against.
The zero did not change. What changed was knowing how ordinary a zero was in the population it came from.
The check that catches all four of these cases is short, and it is worth running before the number goes in a slide.
Ask what else moved during the window you measured. If the answer is anything other than "only the thing I changed," the improvement is a covariance and you should say so out loud rather than let the reader assume otherwise.
Ask what the base rate of your value is in the population it came from. An extreme reading means nothing until you know how common that reading is. A zero drawn from a file that is mostly zeros is a description of the file.
Ask what a fresh sample would say. Not a held-out split of the same collection, which shares every bias the training data has, but a matched sample sourced independently, the way GSM1k was built to mirror GSM8k without inheriting it. If getting one is impractical, that is a real constraint and worth stating. What is not acceptable is spending four years and $62 million without ever noticing that you never had one.
Three questions. None of them require new tooling, and any of them would have raised a hand somewhere in Houston in 2014.
Sources
task_penetration.csv (task level) and job_exposure.csv (occupation level). The task-level and occupation-level counts given above are our own counts of those two files as accessed on July 27, 2026, not figures published by the index.Related reading from us