← Back to blog

The Epistemology of Surrender: When Should a Rational Agent Abandon a Problem?

The Zeigarnik effect does not replicate. Its neglected sibling, the Ovsiankina effect, does, and it is the one that explains a hundred and forty-one years of picking the shovel back up.

Published September 2026 · 13 min read · decision-making / stopping rules / replication / ciphers


To read the only Beale cipher anyone has ever read, you have to count wrong.

The second of the three ciphers is a list of 763 numbers. Each number points to a word in the Declaration of Independence, and the first letter of that word is the letter you want. That much was worked out in the nineteenth century, and the result is a paragraph in plain English about gold and silver buried in Bedford County, Virginia: "the first deposit consisted of ten hundred and fourteen pounds of gold, and thirty-eight hundred and twelve pounds of silver, deposited Nov. eighteen nineteen. The second was made Dec. eighteen twenty-one." But the decoding only works if you number the Declaration the way the original encoder did, and the encoder made mistakes. Wikipedia's technical section lists five modifications to the text that the reader has to reproduce, words to insert and remove, before the numbers line up. The key is not so much the Declaration as one particular person's miscounting of it, and the message is legible only to someone willing to be wrong in the same places.

That is the ground truth of the whole affair, and it is worth holding in mind before anything else, because everything that follows is built on it: the one solid fact in the story is itself strange and partial, and it has been enough to keep people digging for a hundred and forty-one years.

The answers that did not stop anyone

The story arrived in 1885 in a fifty-cent pamphlet called The Beale Papers, published in Lynchburg, Virginia by a man named James B. Ward. It told of a Thomas J. Beale who had buried treasure in 1819 and 1821 and left three ciphers with an innkeeper: one giving the location, one describing the contents, one naming the heirs. The second was the one that read. The first and third, 520 and 618 numbers respectively, never have.

The remarkable part is how many good answers to the underlying questions accumulated, and how little difference they made.

In January 1970 Carl Hammer, director of computer sciences at Sperry Univac, ran the ciphers through a UNIVAC 1108 and found that the numbers were not random. He read this as evidence of a genuine message and pronounced ciphers one and three "for real." He helped found the Beale Cypher Association, whose symposia drew people from the intelligence community, and in 1979 he told the Washington Post that the effort "has engaged at least 10 percent of the best cryptanalytic minds in the country." Note what the computer had actually established: structure. Whether structure meant a message or a fabrication was a separate question, and the answer to it came ten years later.

In 1980 James Gillogly published a short paper in Cryptologia with the title "The Beale Cipher: A Dissenting Opinion." He had done the obvious experiment: decode cipher one with the same Declaration key that unlocks cipher two. The output is noise, except for one stretch in the middle where the letters run almost alphabetically, a b f d e f g h i i j k l m m n o h p p as Nick Pelling later transcribed it. No English plaintext produces that. Gillogly put the odds of it arising by chance in the trillions to one, and offered two readings. Either the message was hidden under a second layer of encryption, or the encoder had been picking numbers more or less at random, grown bored, and started walking down the numbered Declaration in order. The structure Hammer's computer had detected was, on the second reading, the signature of someone making it up. Hammer accepted the finding and rejected the conclusion.

In July 1982 Joe Nickell, in the Virginia Magazine of History and Biography, went at the documents rather than the numbers. The letters supposedly written by Beale in the 1820s used words like "stampeding" that did not enter common use until the 1840s, and their punctuation, grammar and vocabulary matched the pamphlet's narrator closely enough that Nickell concluded the whole thing was written by Ward, the man who sold it.

And in May 2024 Richard Wassmer posted a paper to the IACR ePrint archive with a title that does not hedge: "Beale Cipher 1 and Cipher 3: Numbers With No Messages." Rather than sampling statistics, he shows what the abstract calls "a high correlation between locations of certain numbers in the ciphers with locations in the written text" of the 1885 pamphlet itself, and argues that the two unsolved ciphers were built to have "no intelligible sentences," as part of an elaborate game.

So by the middle of 2024 the three questions a rational person might ask had each been answered as well as such questions get answered. Did Thomas J. Beale write these letters? Probably not; Ward did. Is there a message in ciphers one and three? Almost certainly not; there is an alphabet and a lot of noise. Is there treasure? There is no reason left to think so. None of these is a proof. A negative of the form "no key exists" is hard to establish absolutely, and the honest statement is that this is a strong scholarly consensus rather than a theorem. But that crack, the gap between consensus and proof, is where the diggers live. For more than a century, as the Wikipedia article puts it, people have been arrested in Bedford County for trespassing and unauthorised digging. In 1983 a woman from Pennsylvania rented a backhoe, excavated a church graveyard and unearthed human remains. In the early 1990s a church group from the same state dug pits in the Jefferson National Forest on federal holidays until they were caught and made to fill them in.

What would it have taken?

The question is what it would have taken to stop.

The tempting answer is more evidence, and the record says no. The evidence arrived in 1980, in 1982 and in 2024, from three independent directions, and the digging did not slow. Whatever kept people returning to Bedford County was not a shortage of information, and a stopping rule that says "gather more evidence, then decide" was never going to fire, because no amount of evidence is the same thing as a decision to stop.

Nor is the answer that the diggers were irrational, and I want to be careful here, because it is the cheap move and it is wrong. Under uncertainty about a negative, with a large claimed payoff and a low cost per attempt, continued search can be perfectly defensible. If the odds are one in a million and the prize is several tons of gold, a weekend with a shovel is not a stupid bet. The target of this essay is unexamined continuation rather than enthusiasm: the search that goes on because it went on yesterday, with no one ever having said what would end it.

To see why that happens you need the mechanism, and the brief names the famous one. It is the wrong one.

The wrong effect, and the right one

The famous mechanism is the Zeigarnik effect: the finding, from Bluma Zeigarnik's 1927 study, that people remember interrupted tasks better than completed ones. The cipher is an interrupted task; interrupted tasks stay in memory; the memory nags; the nagging drives the digging.

In July 2025 Romain Ghibellini and Beat Meier published a meta-analysis in Humanities and Social Sciences Communications with the title "Interruption, recall and resumption: a meta-analysis of the Zeigarnik and Ovsiankina effects." Their abstract is direct: the memory advantage of unfinished tasks "has proven particularly difficult to replicate," and pooling the studies, "we found no memory advantage for unfinished tasks." Pooled across the studies, interrupted tasks account for almost exactly half of what people recall, which is what you would expect if interruption did nothing, and setting Zeigarnik's own 1927 data aside does not change it: the ratio is 0.99 either way, across 38 publications with her figures in and 37 without them. The authors are careful rather than dismissive. They attribute the trouble to situational dependencies, "the experimenter's authority, situational demands of task performance, and task involvement," which were common in a 1927 laboratory and are rare now. Zeigarnik was not wrong about her room. The effect just does not travel.

The same meta-analysis found that its neglected sibling does. In 1928 Maria Ovsiankina, working alongside Zeigarnik under Kurt Lewin, published "Die Wiederaufnahme unterbrochener Handlungen" in Psychologische Forschung: the resumption of interrupted actions. Her finding was that people who are interrupted during a task and later given free time show a strong inclination to go back and finish it. Ghibellini and Meier found "a general tendency to resume tasks," and concluded that "the Ovsiankina effect represents a general tendency, whereas the Zeigarnik effect lacks universal validity."

Read that against Bedford County. The diggers do not remember the Beale ciphers better than they remember solved puzzles. Nobody needs to. They go back. The mechanism that survives scrutiny is not a memory effect but a return effect, and it is the one that explains a hundred and forty-one years of people putting the shovel down and picking it up again. Recall was never the issue. Resumption was.

This also separates the Beale case from the concept we reach for first, sunk cost. Sunk cost explains why you do not abandon something you have invested in. Ovsiankina explains why you come back to something you already walked away from. Most Beale diggers are in the second category. They did quit, usually more than once. Then a documentary aired, or a new key text occurred to them, and they were back in Bedford County. A stopping rule designed against sunk cost catches the person who never stopped. It does nothing for the person who stopped and resumed, which is most of us, most of the time, about the problems we cannot leave alone.

One paragraph on the recursion, because it is evidenced rather than merely neat. The abstract's own phrase is that the Zeigarnik effect "has proven particularly difficult to replicate," and the literature kept trying for the better part of a century before a 2025 meta-analysis weighed all of it and found the pooled effect at chance. A field could not find the thing and could not stop looking for it. The research programme on not letting go is itself an instance of not letting go. I do not think anyone involved was foolish; I think they were subject to the effect that does replicate.

The discipline

If the engine is resumption, then the rule has to be built for an agent who will keep coming back, and that constrains its form more than it first appears.

A kill criterion has to be set before you are attached. Once the problem has you, every judgement about whether it is dead is made by the part of you that is generating the problem, and that part will find one more key text. The Beale record shows this in miniature: Hammer could accept Gillogly's finding and still reject Gillogly's conclusion, because acceptance was a judgement and he was the judge.

A kill criterion has to be a state or a date, never a conviction. "When I am convinced it is hopeless" is not a criterion; it is a promise to consult your future self, and your future self is the one who resumed. "If no independent method produces a candidate plaintext by this date, the work stops" is a criterion. So is "if the statistical case against authenticity crosses this threshold, the work stops." Both name, in advance, the observation that ends things, and both can be executed by a calendar rather than by appetite.

And a kill criterion has to be executed, not reconsidered. The moment of execution is the moment the resumption tendency is strongest, because the task is by definition unfinished. That is why pre-commitment is not discipline theatre here. It is the only form of criterion the mechanism cannot route around, because it does not ask the returning mind for permission.

We run research bets in our own work, and the honest sentence of self-application is that every one of them now carries a date and a state written down before the first experiment, because the alternative is a fleet of agents that each behave like Bedford County. Our earlier piece on hospice for legacy systems is about the other half of this: how to wind something down once the decision to stop has been made. This essay is about making the decision, which is the harder half, because nothing in the mechanism wants you to.

The honest asymmetry

I said the diggers were not irrational, and the argument needs its strongest counter-case, which is that sometimes the digger is right and the field is wrong.

On 8 November 1969 the Zodiac killer mailed a 340-character cryptogram to the San Francisco Chronicle. His earlier 408-character cipher had been cracked within days by a Salinas couple, neither of them a cryptologist. The 340 resisted everyone. It resisted for 51 years, until 5 December 2020, when David Oranchak, an American software engineer, Sam Blake, an Australian mathematician, and Jarl Van Eycke, a Belgian programmer, ran roughly 650,000 candidate transposition schemes through Van Eycke's AZdecrypt and found the one that read. The FBI verified the solution. It contained nothing useful about the killer's identity, which is beside the point. Three people who kept returning to a fifty-year-old problem in their spare time were right to.

So the question is not whether persistence ever pays. It plainly does. The question is what distinguished the Zodiac 340 from Beale ciphers one and three, and the answer is not effort or faith. It is that the 340 belonged to a class of problem where a message was known to exist: its author had produced a readable cipher months earlier, and nobody serious doubted that the second one encoded something. The uncertainty was about method, not existence. Beale's first and third ciphers had the opposite profile. Every independent line of evidence pointed at existence, and the answer to the existence question was no.

That distinction is available in advance, and it is the thing a kill criterion should be written against. Before the attachment forms, ask which kind of uncertainty you are holding. If the uncertainty is about method, set a budget and a date and keep going until they run out, because the 340 is real and the people who solved it earned it. If the uncertainty is about existence, name the observation that would settle it, and let the calendar, not your conviction, decide when you have seen it. A confident judgement made later, by someone who has been thinking about the problem for fifteen years, is worth less than a criterion written by the same person before they had thought about it at all. The later self does not know less. It is the one who keeps coming back.


Sourcing notes: the decoded text of cipher two, the five key modifications, the 1819 and 1821 deposits and the "arrested for trespassing" line are from the Wikipedia article on the Beale ciphers, which cites the pamphlet; the counts of 520 and 618 numbers in ciphers one and three are from the ciphertexts; Hammer's 1970 analysis, his "for real" verdict, the 1979 Washington Post quotation and the 1983 and 1990s digging incidents are from Lucas Reilly's Mental Floss account; the alphabetical string is transcribed as Nick Pelling gives it, and Gillogly's odds are stated loosely because two secondary sources render the figure a factor of ten apart; the meta-analysis is quoted from its abstract, verified verbatim through the Semantic Scholar record, and the "almost exactly half" figure from a secondary summary, since the full text sits behind an authentication wall from here, so the specific ratio and percentage reported in the research for this piece are not printed; Wassmer 2024 is quoted from the ePrint abstract page; no dollar value is given for the treasure, and no headcount or hour count is claimed for the diggers, because none exists.

Sources

A criterion written before the attachment forms is worth more than a judgement made after it.

That only works if the criterion, and the date it was written, still exist when the returning mind arrives to argue with them. Chain of Consciousness is where a decision and its basis get written down at the moment they are made, tamper-evident, so the stopping rule cannot be quietly revised by the part of you that wants one more key text.

Hosted Chain of Consciousness

pip install chain-of-consciousness  ·  npm install chain-of-consciousness

Or the whole stack, provenance and ratings and verification together: pip install agent-trust-stack / npm install agent-trust-stack.