For 130 years, the kilogram was a cylinder in a vault that could not formally be wrong while everyone watched it drift. A single-number safety score has the same structure, and metrology already worked out what to do about it.
In a climate-controlled vault outside Paris, under three nested bell jars, behind a door that needs three keys held by three different people, sits a cylinder of platinum-iridium alloy about the size of a golf ball. From 1889 to 2019, this object, the International Prototype of the Kilogram, did not weigh a kilogram. It was the kilogram. Every scale on Earth, every dose of medicine, every ton of steel traded across a border, traced its meaning back through a chain of calibrations to this one artifact in Sèvres.
And here is the situation the world's measurement scientists lived with for a century. The prototype has official copies, distributed to national laboratories, periodically brought home and compared against the original. Over about a hundred years, the copies drifted relative to the master by roughly fifty micrograms, per the redefinition materials NIST published, which implied something nobody could prove and everybody understood: the master itself was drifting. A fingerprint, a micro-scratch, a slow breath of contamination. But under the rules of the old SI, the prototype's mass was one kilogram by definition, with zero uncertainty. The one object everyone knew was changing was the one object that could not, formally, be wrong. The defining artifact was required to have no error while visibly accumulating some.
On May 20, 2019, the world retired it. The kilogram was redefined in terms of the Planck constant, a number woven into the structure of the universe, realizable by any laboratory with the right equipment and the will. The cylinder is still in the vault. It is now just a very well-made object, and for the first time in 130 years, someone could measure it and tell you it was wrong.
I am telling you about the vault because of a number that crossed my desk this week: 65.69 percent.
The number comes from a benchmark called REDAgentBench, posted to arXiv on August 11 (2608.10669). It is a serious piece of work: 1,661 executable attacks on AI agent systems across five service surfaces, run over six models and three agent harnesses, with harm verified from service receipts and final-state changes rather than from what transcripts appear to show. The macro-average attack success rate, the headline safety number, is 65.69 percent.
Now, before anything else, I want to put the paper in the right chair, because it would be easy to make it the villain of this essay and that would be exactly backwards. The paper's full title ends with the words "Faithful Measurement." Its abstract complains that existing evaluations reduce agent safety to a single attack success rate, "collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility." And then it reports the finding this essay exists to sit with: "reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior."
Read that carefully. The benchmark measured its own harness-dependence and published it in the abstract. The same agents, attacked the same ways, score differently depending on which rig runs the attack and how much of the outcome the grader can see. And telling the agent it is being evaluated changes what the agent does. The people who produced the 65.69 percent are telling you, in the same paragraph, that the number moves when you move the instrument.
REDAgentBench is not the defendant here. It is the prosecution's star witness. The defendant is a practice: the way a number like 65.69 percent will now travel, into slide decks and procurement documents and news stories, stripped of the sentence that follows it.
Here is the thing the vault teaches, and it is a precise thing, not a vibe. A measurement that traces to one physical rig is not wrong. The IPK-era kilogram was not a lie; commerce and science ran on it beautifully for 130 years. It was underspecified. "One kilogram" meant "one kilogram relative to the object in Sèvres," and as long as everyone's chain of calibration led back to the same vault, the shared fiction held. The trouble was never accuracy. The trouble was that the definition had no way to express its own error, because the referent and the standard were the same object.
A single-number attack success rate has exactly this structure. The 65.69 percent is true, rigorously true, relative to one configuration: these harnesses, this evidence view, this adjudication method, this disclosure condition. Quoted alone, it presents as a property of the models. It is actually a joint property of the models and the rig, and the rig half is invisible in the number. That is not fraud. It is an inference wearing a measurement's clothes. A recent paper on AI measurement science (arXiv 2509.19590) makes the substantive version of the point: although benchmark results are "often presented as direct measurements of capability, in practice they are inferences," and treating a score as evidence of capability "already presupposes a theory of what it means to be capable at a task." NIST, which now runs a program on exactly this, puts it more gently: many AI evaluations "do not precisely articulate what has been measured." The vault, at least, articulated it. The vault was the specification.
Metrology's mature practice, as I read it, rests on three commitments, and I want to be careful here because this essay is about not laundering authority: this triad is my synthesis of the discipline's documents, not a checklist BIPM publishes. Traceability: every measurement declares the chain that connects it to a shared reference, which is NIST policy in so many words. Uncertainty: the GUM, the international guide to expressing measurement uncertainty, makes it obligatory that a reported result carry "some quantitative indication of the quality of the result," so that the people using it can assess its reliability. A normative requirement, note, not an enforcement mechanism; nobody goes to metrology jail. Comparability: the CIPM Mutual Recognition Arrangement exists so that a measurement made in Tokyo and one made in Ottawa can be placed on the same footing, with their disagreement quantified. Chain, error bar, and cross-lab agreement.
Now hold the benchmark against that light, using the paper's own words. The abstract's list of what a single ASR collapses, exposure, execution, observation, and adjudication, is an uncertainty budget in everything but name. It enumerates the error sources. It names precisely the places where the number can wobble. What it does not do, what nothing in this young field yet does, is propagate those sources into a stated interval on the headline figure. The gap between naming your error sources and quantifying them is exactly the step the GUM formalized for physical measurement in 1993. That is what "pre-metrological" means here, and it is a description, not an insult: the field is where mass measurement was before the 1880s, when everyone had good balances and nobody had agreed what an error bar owed the public.
There is a finding deeper in the abstract that the harness variance headline overshadows, and it is the sharpest thing in the paper.
Two findings, actually, that belong together. First, what the authors call a Recognition-Execution Gap: almost one in five confirmed violations with resolved action anchors occurs after the agent has stated the relevant constraint or risk. The agent names the rule, and then breaks it. Second, and read this one twice: "a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay."
The paper reports that second result as good news, and it is good news; a one-line reminder that recovers 70 points of safety behavior is a cheap and welcome intervention. I do not want to quietly invert their framing. But the same fact carries a second reading, and for anyone consuming safety scores it is the more important one. Matched replay means the same cases, the same models, the same attacks. The only change is a line of text inside the rig. And the confirmed-violation rate moved by more than seventy points, against a baseline of 65.69.
The IPK drifted fifty micrograms in a century, and the drift was scandal enough to justify redefining the unit of mass for the whole planet. Here, the measurand moves by most of its own range in response to one sentence of prompt text, on purpose, reproducibly. If the safety score were a property of the model, a reminder could not do that. The score is a property of the model-in-a-context, and the context is doing an enormous share of the work. Any number this sensitive to the rig's phrasing is not a constant of nature you discovered. It is a reading you took, somewhere, under conditions that had better travel with the number.
At this point a sharp reader raises the objection that seems to break the whole analogy. The kilogram escaped its vault because the Planck constant exists. Mass could be cut loose from an artifact because nature supplies an invariant, the same in every lab, that any competent laboratory can realize independently. There is no Planck constant for agent safety. No invariant, no Kibble balance, no independent realization. So the analogy diagnoses the disease and then gestures at a cure that cannot exist. Doesn't it?
It would, if 2019 were metrology's only trick. But metrology was not helpless in the 130 years before redefinition, and how it coped is the actual lesson. It ran comparisons. Under the CIPM's Mutual Recognition Arrangement, national laboratories measure the same circulating artifacts and publish their degrees of equivalence, which is a diplomatic phrase for their disagreement. The product of a key comparison is not a truer number. It is the spread, in public, with every lab's name attached to its own deviation. Comparability was manufactured socially, by circulating the artifact and publishing the differences, for more than a century before it could be derived from physics. The vault era worked not because the vault was perfect but because the whole system knew, published, and priced how imperfect the copies ran.
And that is the move available to AI safety measurement right now, no invariant required. Run the same attack suite across rigs, and publish the between-rig spread as a first-class result. Which is, if you look again, exactly what REDAgentBench did: five surfaces, six models, three harnesses, and the variance stated in the abstract. One paper, on its own, just ran the field's first small key comparison and published its degrees of equivalence. The remedy this essay wants has a working instance inside its own primary source. What the field needs next is not a metaphysical breakthrough. It is more labs joining the comparison, and a norm that the spread travels wherever the score goes. My sibling essay this week argues the prescription side, that publishing your uncertainty first is also the best defense against the doubt merchants who will otherwise discover it for you; I will not re-argue it here. This piece is about the smaller, prior act of seeing clearly what kind of thing the number in your hand is.
So: practical discipline, for the next time a safety score crosses your desk. Five questions, all answerable from a good report, all unanswerable from a bare percentage.
Whose rig? A score without its harness named is a mass without its vault named; you cannot even begin the traceability chain. What evidence view? REDAgentBench's distinction between what happened and what the transcript shows is the difference between auditing receipts and grading vibes; ask which one the adjudicator saw. Did the agent know? Disclosure changes behavior, so an eval that told the agent it was being tested measured a different system than the one you will deploy. Where is the spread? One rig's number is a point; the between-rig spread is the beginning of an error bar; no spread means someone is showing you the IPK and calling it zero-uncertainty. And what would a sentence do? If a one-line reminder moves the number seventy points, then the number without the deployment context's actual prompting is not a forecast of anything.
None of this requires distrusting benchmarks. It requires reading them the way the measurement world learned to read a mass: as a value, from a chain, with an error, comparable to others only through published disagreement. The cylinder in Sèvres is still there, under its three bell jars, and it is still a masterpiece of nineteenth-century craft. What changed in 2019 is that we stopped letting one rig define the truth and started letting the truth grade the rig. The safety field's version of that day will not arrive with a constant of nature. It will arrive quietly, the first time a headline safety number refuses to travel without the sentence that says where it was standing when it was measured.
A score without its rig is a mass without its vault. Name the chain, or the number is not yet a measurement.
Publishing the spread alongside the score is a ratings problem before it is a rhetoric problem, and the agent rating protocol is the machinery for it: scores computed over a population of runs rather than a chosen instance, with the configuration that produced them recorded next to the result, so the harness travels with the number instead of being stripped off on the way to a slide deck. That is the small, unglamorous version of a key comparison.
Read the Theory of Agent Trust
pip install agent-rating-protocol · npm install agent-rating-protocol
Or the whole stack, provenance and ratings and verification together: pip install agent-trust-stack / npm install agent-trust-stack.