The numerator is mandatory and the denominator is voluntary. Four standard objections to Waymo's safety claim, and what happens when you actually check them.
Here is an asymmetry worth a full minute of your attention. When a Waymo vehicle crashes in the United States, the company must report it: NHTSA's Standing General Order compels crash disclosure from every operator of an automated driving system. But when a Waymo vehicle simply drives a mile, no law anywhere requires anyone to count it. The company's own researchers say so in print: there are no universal state or federal requirements to report automated-driving mileage, so Waymo voluntarily publishes its miles by deployment region.
Every safety rate is a fraction. For this entire literature, the numerator is mandatory and the denominator is voluntary.
That fact does not make the numbers wrong. It makes them the kind of numbers you have to read rather than repeat, and reading them properly is one of the best available exercises for anyone who evaluates systems for a living. Because Waymo's safety claim is, structurally, a benchmark claim. And the way it holds up under the four questions every benchmark reviewer should ask is not the way you would guess.
The load-bearing document is a peer-reviewed paper: Kusano, Scanlon, Chen, McMurry, Gode and Victor, published in Traffic Injury Prevention in 2025, covering 56.7 million rider-only miles through the end of January 2025. Every author is a Waymo employee. Hold that in mind the whole way through; it is the single most important fact about the source, and the analysis below is worth doing precisely because the paper survives it better than you would expect.
The headline findings: across eleven crash-type groups, compared against human benchmarks built from state mileage and police crash data in California, Arizona and Texas, the fleet shows large reductions in injury-involved crashes. Vehicle-to-vehicle intersection crashes with any injury reported: a 96 percent reduction, with a confidence interval of 87 to 99. Airbag-deployment crashes at intersections: 91 percent down. Statistically significant reductions for cyclists, motorcyclists, pedestrians. And in none of the eleven groups, not one, a statistically significant result in the wrong direction.
A skilled reader now reaches for the denominators. Who is in the human comparison group? Is the driving environment matched or merely disclosed? Whose instrument catches more events? Do the severity scales even line up? These are the right questions. They are the questions that dissolve most vendor benchmarks on contact.
What happens next is the reason this essay exists.
Take the environment first. The obvious objection: Waymo's record comes from geofenced surface streets, while the human numbers include freeways, so the comparison flatters the robot. Except the paper matches the restriction on both sides, by construction: the human benchmark is filtered to surface streets too, using the federal highway function classifications, with the exclusions documented state by state. The freeway objection is not an unmatched denominator. What remains is a generalization limit, and a real one: the result tells you nothing about freeway driving. But "this doesn't generalize past its envelope" is a much weaker charge than "this comparison is rigged," and a reader who conflates them will lose to anyone who has actually read the methods section.
Take reporting asymmetry next, because it is the strongest objection in the set, and the paper's handling of it is the tell that you are reading serious work. A Waymo vehicle's telemetry reports every scrape automatically; human fender-benders reach police data only when someone bothers. The authors state this against themselves, assume little underreporting in their own data, and then do the honest thing: they drop the lowest-severity comparison entirely rather than publish a category they would win by instrument sensitivity. On the injury comparison, they adjust the human benchmark upward to account for the estimated 33 percent of injury crashes that never reach police reports, attributing the estimate to Blincoe and colleagues, and they run a sensitivity analysis showing the conclusion does not hinge on that number. Where they could not justify an adjustment, on the most serious injuries, they note the omission makes the comparison conservative.
Notice the direction of every one of those moves. Each named adjustment either shrinks the measured advantage or deletes a comparison the sponsor would have won. Adjustments that only ever help the sponsor are the classic tell of a cooked benchmark. Here the arrow points the other way, which is evidence of method. Not virtue, method. The distinction matters and the essay's job is not to canonize anyone.
The third objection is the one everybody makes first, and what the evidence does to it is my favorite thing in this whole story.
The human crash rate includes drunk drivers, exhausted drivers, texting drivers, unbelted drivers. So the benchmark human is worse than the median attentive one, and beating that population is a low bar. Directionally reasonable. Every skeptic says it. I would have said it.
Now look at the one axis where somebody actually measured it. Impaired and risky driving concentrates at night. And Waymo's fleet is over-exposed to the night. Per Chen and colleagues' 2024 analysis, as reported in the Kusano paper: in San Francisco the fleet drove 41 percent of its miles in the evening and overnight hours against 24 percent for human traffic, and only 16 percent in the daytime window against a human 40 percent. Correcting the benchmark for time of day therefore makes the human comparison worse, not better: the adjusted benchmark comes out 1.05 times higher for any-injury outcomes and 1.16 times higher for airbag deployments. The unadjusted comparison understates the fleet's advantage.
Read that twice, because the shape of it transfers everywhere. The intuitive correction, applied without measuring exposure, would have moved the number in the wrong direction. The fleet drives disproportionately in exactly the hours the drunk-driver objection is about. A critic who adjusts for the risky population without adjusting for exposure to the risky hours has made the comparison less accurate while feeling more rigorous.
Scope it honestly, as the paper does: this adjustment exists for one city, extrapolated from a single traffic study, and it was not implemented in the headline analysis because equivalent data does not exist elsewhere. So the defensible statement is narrow: on the one axis where the intuition has been tested, it inverted. Everywhere else it remains an intuition.
So the four standard objections go: matched, handled, handled, inverted-where-measured. That is not the end of the reading. It is where the interesting residue starts, most of it stated by the authors themselves.
Fault is not measured. The study counts crash involvement, not contribution; a vehicle that gets rear-ended more often can look worse while driving better, and the paper concedes even insurance claims are an imperfect proxy for responsibility.
The severity ladders are not the same ladder. The federal reporting order's "serious injury" turns on emergency treatment, a much lower threshold, the paper notes, than the hospitalization standard behind police-reported incapacitating injuries. The same word, two different thresholds.
The matching is unbounded, and they say so: there is, in the authors' words, an endless list of possible factors to potentially account for. That sentence is the honest description of every matched comparison ever constructed. You cannot match on everything. You can only publish what you matched on, and let the reader judge what you left out.
And then there is the one the paper does not say. Search all fifty-eight pages of the preprint for the word "weather" and you get zero hits. Snow, zero. Fog, zero. The only match for "rain" is inside the word "restrained." Weather is not matched, not adjusted, not sensitivity-tested, and not listed in the limitations, across deployments in Phoenix, San Francisco and Texas, three of the mildest driving climates in America, compared against human mileage that includes every storm those states had. To be precise about what this is: an absence in the published methodology, verified by searching the text, not a concealment, and not a claim about what the fleet chooses to drive in, which I have not verified. But after watching this methodology answer four obvious objections in print, the strongest surviving objection turns out to be the one nobody wrote down. That, too, is a lesson that travels: audit the limitations section for what it omits, because the authors' own list of weaknesses is also a curated dataset.
The AV industry already has a producer-side standard for this: the RAVE checklist, assembled by researchers from industry, insurance and academia, which this paper explicitly conforms to. That is a checklist for people running retrospective safety studies. What does not exist, and what most of us actually need, is the consumer-side version, for reading any safety-or-performance claim built on a comparison. Here is mine, earned from this case:
Who is in the comparison group, and does the claimant control membership? Waymo does not pick the human drivers, but it picks where and when its fleet drives.
Is the envelope matched or merely disclosed? Surface streets on both sides is matching. A footnote is disclosure.
Whose instrument is more sensitive? If one side's telemetry catches everything and the other side's data needs a police report, low-severity comparisons are noise, and the honest move is to drop them, as this paper did.
Are the severity scales the same scale? Look for the same word meaning two thresholds.
Who mandates the numerator, and who mandates the denominator? If the answers differ, part of the rate is voluntary.
Which way did the published adjustments move the result? All in the sponsor's favor is a tell. Mostly against, with a sensitivity analysis attached, is what method looks like.
What is absent? Search the document for the confounder that worries you most. Zero hits is a finding.
Does the quote generalize past the envelope? "Safer on surface streets in three states through January 2025" is what was measured. "Safer than humans" is what gets repeated.
Strip away the cars, and here is what this Waymo team did: they disaggregated their headline result into eleven slices, tested each one, and published the null results alongside the wins, including the explicit statement that no slice showed a statistically significant regression. Then they published their adjustments, their sensitivity analyses, and a limitations list, under an external reporting checklist, with their names on it.
Now think about the last model evaluation you shipped or accepted. One aggregate number, probably. No per-slice breakdown, no adjustment ledger, no statement of which corrections were considered and rejected, no sensitivity analysis on the labels. The comparison every eval implies, against a baseline whose data was gathered by someone else's instrument under someone else's incentives, gets one line if it gets anything.
The uncomfortable conclusion of reading Waymo's safety numbers closely is not that they are weaker than advertised. Within their envelope they are strong, the envelope is written down, and the writing down is more than most benchmarks manage. The uncomfortable conclusion is that a car company's safety team is currently ahead of most of the software industry at reporting comparative claims, and the gap is not talent. It is that someone made them write down the denominator.
Nobody is making you. Do it anyway. Next time a number arrives wearing the word "safer" or "better," read the fraction before you repeat it: who counted the top, who chose the bottom, and what word appears zero times in the fifty-eight pages underneath.
Sources: Kusano, Scanlon, Chen, McMurry, Gode and Victor, "Comparison of Waymo Rider-Only Crash Rates by Crash Type to Human Benchmarks at 56.7 Million Miles," Traffic Injury Prevention 26(sup1), 2025; full-text preprint at arXiv:2505.01515, from which all quotations and figures are drawn, including its characterizations of Chen et al. (2024) on time-of-day exposure, Blincoe et al. (2023) on underreporting, and the RAVE checklist (Scanlon et al. 2024), none of which were independently retrieved for this essay; NHTSA Standing General Order on ADS crash reporting, as described in the same paper. All study authors are Waymo employees; the paper is cited as the peer-reviewed methodology under review here, and this essay does not adjudicate the separate disputes between Waymo and its independent critics (Cummings and Bauchwitz 2024; Chen and Shladover 2024), whose papers were not read for this piece.
Somebody made them write down the denominator
The reason this paper survives its own audit is that its comparison is written down: who is in the benchmark, which way each adjustment moved the result, and what was dropped rather than won on instrument sensitivity. Agent ratings have none of that by default. A score arrives, the population it was measured against does not, and the fraction is unreadable. Agent Rating Protocol is a rating format that carries its own denominator: what was evaluated, against whom, and on whose instrument, so a claim about one agent being better than another is a fraction you can read rather than a number you have to trust.
pip install agent-rating-protocol · npm install agent-rating-protocol