← Back to blog

The Wiki Says B Minus A Minus Two C

A department measures water the long way for eleven years. The research says the exception rate at that dosage is zero.

Published July 2026 · 11 min read · cognitive bias / process design / onboarding / runbooks


The onboarding wiki has one page for measuring water, and it has been the same page for eleven years.

Standard Volumetric Retrieval Procedure (SVRP)

You will be issued three jars: A (21 units), B (127 units), C (3 units).

To obtain any target volume, fill B. Pour off A once. Pour off C twice.

Confirm the remainder. File the confirmation.

Example: target 100. Fill B (127). Pour A (−21). Pour C twice (−6). Remainder: 100. ✅

The new hire, Marta, reads this on her first morning and thinks: fine. Odd, but fine. It works. She runs the procedure on her first ticket and the number comes out exactly right. She runs it on her second and third. It comes out exactly right both times. By lunch on day two she has stopped reading the jar labels, because the shape of the thing is in her hands now: big jar, small jar once, tiny jar twice.

On day three she gets a ticket where the jars are A = 23, B = 49, C = 3, and the target is 20. She runs the procedure. Forty-nine minus twenty-three minus six. Twenty. Files the confirmation.

Then she looks at it again, because twenty and twenty-three are three apart, and there is a three-unit jar sitting right there.

She fills A. She pours off C. Twenty.

One step. She checks it twice, then walks it over to Dev, who has been measuring water here since before the merger.

This is where our story becomes a documented experimental result, and stays one for the rest of the essay.

The department is a laboratory that has been running since 1942

Abraham Luchins ran exactly this setup on human subjects in 1942: a set of training problems solvable only by the long formula, then test problems solvable by both the long formula and a trivially short one. Our readers have met the water jars here before, in a piece on how breakthroughs happen, so I will not present them as a reveal. What that piece did not carry are the numbers underneath, and the numbers are where the satire lives.

They are worth stating with their denominators, because the version of this study that circulates online has lost them. The counts below come from Marcel Binz and Eric Schulz's 2021 reconstruction, which transcribed Luchins's own tables and re-analysed them. Luchins published percentages; the counts are theirs.

The base result. Subjects who had been trained on the long method used it on the first pair of test problems between 77% and 82% of the time, even though a two-step solution was available and visible.

The control group is the fact that makes it a phenomenon and not a personality flaw. Subjects who got no training problems at all never used the long formula. Not rarely. Never. Nobody looks at three jars and spontaneously invents "fill the big one, pour off the small one once and the tiny one twice." The method does not arise. It is installed.

So when Dev tells Marta that the procedure exists for a reason, Dev is right in a way Dev cannot articulate. The reason is that somebody trained him on it.

Day four: alignment re-training

Marta is invited to a meeting. It is a kind meeting. Nobody is angry.

Dev explains that the one-jar approach is clever, and that clever is not the same as correct. He explains that the SVRP is proven, that it handles cases her shortcut does not, and that it scales. He says "that's not really how we do things here," and he says it warmly, because he likes her.

Twice during the meeting, mid-sentence, his eyes go to the big jar.

That detail is not a novelist's flourish. Bilalić, McLeod and Gobet demonstrated it with chess experts and an eye-tracker in 2008, and the cleanest restatement of what they found appears in Sheridan and Reingold's 2013 study of the same effect (their Einstellung paper, not their earlier work on expert visual span):

"Replicating prior findings, the chess experts initially discovered the familiar but longer solution (i.e., checkmate in five moves), but failed to find the shortest solution (i.e., checkmate in three moves)."

And then the sentence that should be pinned above every runbook:

"the chess experts continued to look at the chess squares associated with the familiar solution, even though they reported that they were searching for alternative solutions."

The experts were not lying about searching. They were searching. Their eyes went back to the familiar squares anyway. Dev is not pretending to consider the one-jar method. He is considering it, in a room, on a board where the good move is already lit up.

And here is the part of the meeting that matters most, which nobody in it notices: the SVRP has to actually work for any of this to happen. Sheridan and Reingold tested 34 players, 17 of them experts averaging 2223 rating, and found the boundary condition precisely: when the familiar move was merely suboptimal, both experts and novices kept staring at it. When the familiar move was an outright blunder, all of the experts avoided it.

A method that fails gets abandoned. A method that succeeds slightly worse than necessary is nearly immortal. If the SVRP got the wrong answer, Dev would have dropped it a decade ago. It gets the right answer. That is the trap. The department is not held hostage by a bad process; it is held by a good one that stopped being the best one and never announced it.

Day forty: the ticket that cannot be done

The ticket arrives at 4:40pm on a Friday, marked urgent, and the jars are A = 28, B = 76, C = 3, target 25.

Run the SVRP. Seventy-six, minus twenty-eight, minus six. Forty-two.

Run it again. Forty-two.

Somebody suggests running it more carefully.

What happens next is the most replicable finding in this literature, and it circulates in a mangled form worth correcting, since a piece about unexamined inherited formulas should not pass one along. You will see it quoted as a 97% failure rate, sometimes attributed to being observed. The 97 is real; both labels are wrong. In the reconstruction's tables, subjects under time pressure used the long method on 95 of 98 test problems, up from 126 of 153 without it. Roughly 82% becomes roughly 97%.

The right word for it is reversion. Those subjects all got the correct answer; they got it the slow way. And the trigger was a deadline, not an audience.

Everyone knows this feeling and almost nobody has seen it quantified. Under deadline, you do not get more inventive. You get more like yourself. The department at 4:40 on a Friday is a department that has lost the ability to consider a second method precisely when it most needs one, and the loss is not a character failing. It is a measured effect with a slope.

The team escalates. Dev asks for a bridge call. Someone starts a document titled SVRP Edge Cases (WIP).

The intern, who started on Tuesday and has not been trained on anything, is sitting in the corner because nobody has assigned her a desk. She looks at the jars for a few seconds.

She fills A. She pours off C. Twenty-five.

The intern is not a comic device. She is the control group.

This is the best structural fact in the research file and it deserves to be said plainly. In the same passage where Sheridan and Reingold restate the gaze finding, they report the control condition:

"as evidence that the optimal solution was not inherently difficult, a control group of chess experts successfully discovered the optimal solution when they were shown a modified version of the problem that did not contain the familiar solution"

Same experts. Same difficulty. Remove the familiar move from the board and they find the good one immediately. The intern is not smarter than the department; she is a department member with the training removed. Every single person in that room could have poured 28 minus 3, and would have, if they had never been shown the wiki.

Which brings the piece to its real subject. The training that produces this does not have to be bad at all. It only has to be repeated.

In Luchins's dosage condition, subjects who received two training problems used the long method on 79 of 108 test problems, about 73%. Subjects who received ten used it on 112 of 112.

One hundred and twelve out of one hundred and twelve. There is no dissenting minority at ten repetitions. When I write that the senior staff have measured water this way for eleven years and not one of them has questioned it, I am not exaggerating for comic effect; the literature says that at sufficient dosage the exception rate is zero. Eleven years is my invention. The ceiling is not.

The correction we owe our own back catalogue

Our earlier piece described the Einstellung effect as "the cognitive bias where prior successful strategies prevent you from finding better ones." That is the standard framing, it is what most write-ups say, and the 2021 reconstruction argues it is the wrong shape. From Binz and Schulz's abstract:

"We furthermore show that a model of resource-rational decision-making can explain all of the observed effects. This model assumes that people transform prior preferences into a posterior policy to maximize rewards under time constraints. Taken together, our reconstructive and modeling results put the Einstellung effect under the lens of modern-day psychology and show how resource-rational models can explain effects that have historically been seen as deficiencies of human problem-solving."

Read what that does to the satire. Dev is not stupid. Dev is running what our piece on sycophancy as resource-rational behaviour would call the correct policy for an agent with finite search budget and a method that has returned the right answer on every observed trial. Searching for a better approach costs something; the expected saving from re-deriving an already-solved problem is small; the rational move is to reuse. He would be wrong to search under the assumptions he is carrying.

The assumptions are the problem. Not the man.

That reframing is where the piece stops being a joke about a slow colleague and becomes something a manager can use, because it relocates the defect from a person to an unowned job. The SVRP is a cached decision. It was correct when it was cached. Somewhere between then and now, the cache's assumption expired — the jar sizes changed, or the constraint that made the long path necessary quietly lifted — and nothing in the building was responsible for noticing. There is no cache-invalidation step because writing one was never anybody's task. Every character behaves reasonably and the department still walks off a cliff, which is the only kind of satire worth writing.

(This is also, for anyone who read our piece on the streetlight effect, a different animal. That one is about searching where the light is. This one is about solving where the method is: not a bad search, a finished one.)

The debiasing, with prices, because that is the part you can act on

Two interventions were measured. They are not equally good, and the gap between them is the practical content of this entire essay.

Tell people about the trap. In one condition, subjects were given a warning — the substance of "don't be blind" — before the test problems. Long-method use fell to 87 of 153, about 57%, down from roughly 83%.

Change how they were trained. In the alternating condition, where problem types were interleaved rather than run in an unbroken block of the same shape, long-method use fell to 20 of 124, about 16%, against 97 of 121 (roughly 80%) for the blocked training.

A warning cuts the effect by about a third. Interleaved training cuts it by about four fifths.

That is the whole management lesson in two rows, and it says something uncomfortable about how most organisations respond to this problem. The warning is the all-hands slide about avoiding assumptions, the retro action item that says "consider alternatives," the sentence at the top of the wiki reading this is one approach among several. It measurably works, it is cheap, and you should do it.

But it is a third of the available benefit, and it is the intervention we reach for because it costs nothing structural. The other four fifths live in the curriculum. If every onboarding ticket a new hire touches in their first month has the same shape, you are running the blocked condition on purpose, and you will get its result: not people who can't think of the short path, but people who reliably don't, fastest exactly when it matters most.

What to do on Monday

Three things, in ascending order of cost and effect.

Put the warning in. One sentence at the top of the runbook naming the trap: this procedure is general, and most tasks have a shorter path — check for one before running it. A third of the effect for the cost of a sentence is the best trade in this essay.

Interleave the onboarding. Deliberately vary the shape of a new person's first weeks. Do not let them run ten of the same ticket, because ten is the documented dose where the exception rate reaches zero. Mix in tasks the standard procedure handles badly, and mix them in early, while the method is still a suggestion rather than a reflex.

Give cache invalidation an owner. This is the one nobody does. Every well-established procedure encodes a constraint that was true when it was written. Someone has to hold the job of asking, on a schedule, whether that constraint still holds — not whether the procedure works, because it does, that is the trap, but whether the reason for it survives. Put a date and a name on the page. "Reviewed against current tooling, March, by Dev." A runbook without a review owner is a bet that the world stopped changing on the day it was written.

And when someone new asks why you do it the long way, notice what your eyes do while you answer.


Marta gets promoted, incidentally. So does the intern, who is asked to take over the wiki, since she is the one who solved the Friday ticket.

She opens the page template. It has fields for Procedure, Worked Example, and Confirmation Step. She has one step and no formula, and the form wants three sections.

She fills the big jar, pours off the small one once and the tiny one twice, and writes down what happens.


Sources

If your runbook encodes a reason nobody wrote down

The expensive failure in this essay is not the long procedure. It is that the reason for the procedure was never recorded, so no one could check whether the reason still held. Chain of Consciousness keeps the decision trail an agent actually followed — what was decided, and on what basis — which is the record a cache-invalidation review needs to have something to read.

pip install chain-of-consciousness

npm install chain-of-consciousness

If you would rather not run it yourself, the Hosted CoC page covers the managed version.