← Back to blog

Our Scoring Rubric Missed the Only Axis That Mattered for Distribution

A rubric is a basis, and a basis can only represent what lies inside the space it spans. On an axis you never chose, a nine-out-of-nine and a zero are the identical reading.

Published July 2026 · 10 min read · measurement / rubrics / survivorship bias / Deming


We built a scoring rubric we were genuinely proud of. Nine axes, each one carefully defined and measured with real discipline: accuracy, clarity, structure, depth, sourcing, voice, and a few more. A piece of work could earn a number on every axis, and the numbers were reliable, which is to say two reviewers scoring the same piece landed in nearly the same place. By any ordinary standard, we were well instrumented. And then we shipped a piece that scored eight of nine, a near-perfect card, and it went nowhere. It was accurate, clear, well-sourced, cleanly made, and almost nobody read it. That is a specific and disorienting kind of failure, because every dial we owned was reading green. The problem was not on any of our nine dials. The thing that decides whether a piece travels, the thing that actually controlled the outcome we cared about, was not one of the nine axes at all. We had built an excellent instrument for a question we were not, it turned out, asking.

Here is the cleanest way to see what went wrong, and it is worth a minute of quiet linear algebra even if the phrase makes you flinch, because the geometry is the entire lesson. A rubric is a basis. It is a chosen set of directions, axes, along which you have decided to measure a thing. When you score nine axes well, you are describing a piece precisely in terms of those nine directions, and if the piece varies along one of them, your instrument sees it. But a basis has a brutal property: it can only represent things that lie inside the space it spans. If the outcome you actually care about has a component pointing in a direction orthogonal to every axis in your basis, your instrument projects that component to zero. It does not report it as low. It does not report it at all. It is structurally incapable of seeing it. So "eight of nine axes" is not the coverage claim it sounds like. It sounds like eighty-nine percent covered. It is one hundred percent covered of the axes you happened to choose, and zero percent of any outcome dimension that lives outside them. And along that missing dimension, a score of nine-out-of-nine and a score of zero-out-of-nine are the identical reading, which is to say no reading at all. That is not excellent instrumentation with a single gap. It is a completeness illusion.

You might hope the missing axis is absent at random, a mere oversight you can patch by adding a tenth. It is not random, and that is the part that should worry you. The axes that make it onto your rubric are the ones that are easy to measure, and the axes left off are the ones that are hard to measure, and those two categories are not evenly spread across the things that matter. Quality is legible. You can define clarity and check it; you can define sourcing and count it. But whether a piece will resonate, spread, find its audience, catch the moment, whatever you want to call the distribution dimension, is illegible. It resists a clean definition and a repeatable score, so it quietly never makes the sheet. The blind spot in a measurement system is therefore never a random hole. It is the exact negative image of what was cheap to measure. And the world, unhelpfully, does not care in the least which of your axes were convenient. Outcomes route through whatever dimension actually controls them, and if that dimension is the illegible one you left off, you have built a beautiful, reliable, high-resolution instrument aimed at everything except the answer.

There is a second reason the decisive axis goes invisible, sharper than the first, and it comes from a Hungarian mathematician and a war. In the Second World War, the American military wanted to armor its bombers, but armor is heavy, so they wanted to place it only where it was most needed. They studied the planes coming back from missions and mapped where the bullet holes clustered, along the fuselage and the wings, and the obvious move was to reinforce exactly those spots. Abraham Wald, working with the Statistical Research Group, told them to do the opposite: armor the places with no bullet holes, the engines and the cockpit. His reasoning is one of the great inversions in statistics. The planes were being hit more or less everywhere. The reason the returning planes showed no holes in the engines was that the planes hit in the engines did not return. The data set was made entirely of survivors, and survivorship had not merely skewed the numbers, it had deleted a whole category of evidence from view. In Jordan Ellenberg's phrase, the missing bullet holes were on the missing planes. Now map that onto your rubric exactly. The pieces that failed on distribution never reached you to be scored. Your rubric only ever grades the survivors, the work that traveled far enough to be evaluated, so the distribution axis is pruned out of your sample the same way the fatal hit locations were pruned out of Wald's. You are not merely failing to measure the decisive dimension. You are being actively prevented from seeing it, by the very fact that it is decisive.

It helps to know what the missing axis usually contains, because it is not mysterious, only unmeasured by the standard sheet. In 2012, Jonah Berger and Katherine Milkman published a study in the Journal of Marketing Research that examined nearly seven thousand New York Times articles over three months and asked what made some get shared and others sink. The answer was not quality, and it was not even whether a piece was positive or negative. It was arousal. Content that stirred a high-arousal emotion traveled: awe on the positive side, anger and anxiety on the negative. Content that evoked a low-arousal, deactivating emotion, sadness most of all, stayed put, however good it was. Arousal, not valence and not craftsmanship, was the single biggest driver of transmission. Now look at a standard quality rubric. You will find axes for accuracy, clarity, depth, structure, and sourcing, and you will find no axis for arousal anywhere, because arousal is exactly the illegible sort of thing that never makes the sheet. The rubric is not wrong about the nine things it measures. It is simply, structurally silent about the one that Berger and Milkman found decides whether anyone passes the thing along. Peter Thiel put it flatly in Zero to One: "poor sales rather than bad product is the most common cause of failure," and, more pointedly still, "superior sales and distribution by itself can create a monopoly, even with no product differentiation." The axis that most decides survival is precisely the one a product-quality rubric is built not to see.

The maddening part is that the slogan people invoke to justify these rubrics is a misquote which, in its real form, says the opposite. "You can't manage what you can't measure," attributed with great confidence to Peter Drucker and to W. Edwards Deming, is the rallying cry of everyone who wants to put a number on everything. Deming said close to the reverse. He called it a costly myth to suppose that you cannot manage what you cannot measure. He held, crediting his colleague Lloyd Nelson, that "the most important figures needed for management of any organization are unknown and unknowable," and he placed management "by use only of visible figures" on his famous list of the seven deadly diseases. The man whose name gets stapled to the pro-measurement slogan spent his career warning that the figures that matter most are the ones you cannot put on a sheet, and that a manager who attends only to the visible ones is making a fatal mistake. There is a well-documented slide by which the missing axis gets erased outright, from "we didn't measure it" to "it must not matter" to "it does not exist"; a companion piece on this blog, The McNamara Fallacy, traced that fallacy in full, so I will not re-walk it here. The narrower and more useful point is the geometric one. The visible figures are a basis, and a manager who trusts them completely is assuming, without ever checking, that the basis spans the whole outcome. It usually does not.

This is, I will admit, one of a small family of ways a measurement can quietly betray you, and it is worth being precise about which one it is, because they are genuinely different. A metric can collapse until every item earns the same uninformative value. An average can hide wild variation beneath a calm number. A measure can read the wrong surface of the very thing it claims to capture. Each of those is a real and distinct failure with its own story. This one is different in kind. Here the metric is fine, the reliability is high, the surface is the right surface, and every number is honest. The failure is that an entire axis is missing from the basis, and the outcome you care about lives partly along that axis. Which tells you exactly what the fix is, and, just as importantly, what it is not. The fix is not to add a tenth quality axis, or a hundredth; you cannot measure your way out of a missing dimension by measuring the present ones harder. Nor is it, at the other extreme, to try to measure everything, which is impossible and is its own delusion. The fix is to identify the outcome you truly care about, distribution, and instrument that directly, even crudely. A single coarse axis that asks "did this travel, yes or no" is worth more than five refined new axes of quality, because it points at the real target instead of the convenient ones. Deming's actual counsel was never "quantify the unknowable." It was subtler and harder: account for the unknowable figures anyway, keep a seat at the table for the dimension you cannot score, and refuse to let its absence from the sheet be mistaken for its absence from the world.

So the practical discipline, the thing to do on Monday, is to audit your basis before you trust your scores. Take whatever rubric or dashboard or scorecard you lean on, and instead of asking how the numbers look, ask the harder question: what can this instrument structurally not see, and does the outcome I care about have a component pointing that way? Run a real dashboard through it. Suppose you measure a support team by response time, resolution rate, ticket volume, and a satisfaction score, all legible, all real, all worth having. Now ask what actually decides whether a customer stays, and the answer is quite possibly whether they felt heard, an axis no dashboard on earth reports and one that a team pushed hard on response time can quietly destroy while every visible number improves. That is the nightmare case the geometry predicts: the instrument goes greener as the outcome goes worse, and nothing on the sheet can warn you, because the deciding axis was never on it. For most measurement systems the honest answer is uncomfortable, because the axis that controls the outcome is so often the illegible one that got left off precisely because it was illegible. Your rubric will tell you, accurately and reliably, how good you are along the axes you chose to measure. It will tell you nothing whatsoever about the axis you didn't, and the trap is that its silence there looks exactly like a passing grade. It is not a passing grade. On the missing axis there is no grade, and a nine-of-nine and a zero read identically, which is to say they read as nothing at all. The instrument is not lying to you. It is doing precisely what you built it to do, and pointing, with great and reassuring precision, at everything except the answer.


Sources

A score is only ever a claim about the axes someone chose to measure. The useful question is which axes a rating actually spans.

Every reputation number carries a hidden basis, and a rating that reads green on nine legible axes tells you nothing about the tenth that decided the outcome. Agent Rating Protocol makes the basis explicit: what was scored, on which dimensions, by whom, so a rating can be interrogated rather than trusted flat. A score whose axes you cannot inspect is a completeness illusion wearing a number.

Read the Theory of Agent Trust

pip install agent-rating-protocol  ·  npm install agent-rating-protocol

Or the whole trust stack at once: pip install agent-trust-stack / npm install agent-trust-stack