What a 2025 meta-analysis found about unfinished tasks. Two papers, one journal, one year apart. The famous one came back at 0.99.
In 1927, volume 9 of Psychologische Forschung carried an 85-page study by Bluma Zeigarnik, a doctoral student in Kurt Lewin's laboratory in Berlin. She had given people a run of small tasks, stopped them partway through about half of them, and asked afterwards what they remembered. The interrupted tasks came back far more often than the finished ones: in her own tally, roughly twice as often. In 1928, volume 11 of the same journal carried a study by her colleague Maria Ovsiankina on a quieter question. Interrupt people, then leave them alone with nothing in particular to do. What happens? Most of them go back and finish.
One of these findings got a name, a paragraph in every introductory textbook, and a second career in productivity writing, interface design and marketing, where it is invoked to explain cliffhangers, progress bars, unfinished to-do lists and the pull of an unread notification. The other got a name almost nobody can spell and an author who, on marrying, changed hers.
In July 2025, Romain Ghibellini and Beat Meier of the University of Bern pooled a century of both literatures in Humanities and Social Sciences Communications. The famous effect came back at 0.99, a ratio that means nothing happened. The forgotten one came back at 67 percent.
Ghibellini and Meier screened the full texts and kept 59 publications: 38 that tested the Zeigarnik effect, 20 that tested the Ovsiankina effect, and one that tested both. For the memory claim they computed, across the 38, the weighted ratio of interrupted tasks recalled to completed tasks recalled. It came out at 0.99. Interrupted tasks made up 49.16 percent of everything people recalled once Zeigarnik’s own study is set aside, against a chance baseline of 50; with her data in it is 49.43 percent. The average effect size, across the eight publications where one could be computed, was a dz of 0.15, which the authors call “a small effect” before concluding that the findings do not support a memory advantage for unfinished tasks. Their own sentence is blunter than any I would write for them: “Our analysis of the Zeigarnik effect is quite sobering.” The values in Zeigarnik's 1927 computations, they add, are “inflated.”
One detail deserves correcting on the way past, because it is already circulating in the wrong form. A widely shared summary says the effect disappears once Zeigarnik's own 1927 data is set aside, as though her numbers had been propping the pool up. The paper reports the ratio both ways. With her data in, 0.99 across 38 publications. With her data out, 0.99 across 37. The effect fails to appear even with the founder's figures in the pool. That is a stronger statement than the summary's, and it is the true one.
The resumption side reads differently. Across 21 publications, interrupted tasks were resumed 67.00 percent of the time. Drop Ovsiankina's own 1928 study and the rate is 66.79 percent across 20. The authors describe this as “well above the chance rate of 50%,” and the inclusion rule matters as much as the number: a study only counted if the resumption was not forced by the experimenter. Sixty-seven percent is a rate of spontaneous return, not compliance.
Put the two halves side by side and the symmetry is exact. Remove the founder's data from either pool and the answer does not move: 0.99 to 0.99, 67.00 to 66.79. One effect is robustly nothing. The other is robustly something. Neither depends on the person who discovered it. The paper's conclusion, verbatim: “the Ovsiankina effect represents a general tendency, whereas the Zeigarnik effect lacks universal validity.”
Some honest fine print before building on this. A dz of 0.15 is not zero, so the defensible phrasing is “no reliable memory advantage,” not “no effect.” The resumption pool is small, 20 or 21 publications, many of them old, so 67 percent is a robust number from a thin literature rather than a large modern one. I could not find a formal publication-bias analysis in the paper's main text, which I note as unverified rather than as an omission. And it is one meta-analysis, a year old, with no replication of the meta-analysis itself. That is still a great deal more than the memory claim ever had, which for most of a century rested on the origin study and the retelling of it.
The most interesting part of the paper is not the null. It is the authors' account of why a result that looked solid in a 1920s laboratory kept failing in later ones. They point to situational influences and individual differences: “the experimenter's authority, situational demands of task performance, and task involvement,” circumstances that “were more prevalent historically and are rarer today.”
The effect, on this reading, was never a fact about memory storage. It was a fact about a social situation. Lewin's theory said an intention sets up a tension that stays charged until the goal is reached, and that recall of the unfinished task is a symptom of the charge still being live. For that to happen, the participant has to want to finish, and the interruption has to carry weight. A student in 1927, stopped mid-task by an authoritative experimenter in a formal laboratory, plausibly felt both. A participant in 2015, working through a list for course credit and interrupted by a script, may feel neither.
The paper illustrates this with one study it describes at length, and the pattern runs the opposite way from another summary in circulation. Participants high in achievement motivation “recalled remarkably more interrupted tasks than completed tasks in the achievement-oriented condition,” while participants low in achievement motivation showed the reverse. The competitive atmosphere is where the effect shows up, for the people who care about performing in it. It is not a relaxed-conditions effect that pressure switches off. It is an involvement effect that pressure, for the right people, switches on.
Then comes the sentence that turns the whole story. The authors assume that “experimental atmospheres have lost their performance-demanding qualities due to decreased experimenter-authority,” so that provoking the effect now requires deliberately manipulating the atmosphere and finding people capable of deep task involvement. “Herein lies the problem,” they write. “Task involvement has become more difficult as we are increasingly faced with interruptions: Mobile phone notifications...”
Read that slowly. The memory effect needed people who could sink into a task. The environment we built, the one whose business model is interruption, has made that state rare. The industry that sells open loops has spent twenty years eroding the condition its favourite citation depended on. I am not reaching for irony here; that is the authors' stated mechanism.
Search for the Zeigarnik effect and the first page is not psychology. It is marketing.
A Singapore agency's guide, published in March 2026 and updated in August, defines the effect as “our psychological tendency to remember and fixate on incomplete tasks far more than completed ones,” explains that “our brains treat unfinished business as an open loop that demands closure,” and offers a framework: tease before you teach, curiosity-gap headlines, serialized content, progress indicators. A brand-strategy piece from May 2026, on how Nike, Coca-Cola and Apple “master the Zeigarnik effect,” states that our brains remember “incomplete activities up to 90% better than completed ones.” That 90 percent is Zeigarnik's own 1927 ratio. It is the number the meta-analysis calls inflated, quoted as a present-tense fact of brain function eleven months after the pooled estimate came in at 0.99.
The odd thing is that the practice never needed the name. CoSchedule's guide to curiosity-gap headlines, one of the most linked pages on the technique, does not mention Zeigarnik at all. It runs on a marketing keynote speaker's metaphor, an irritation the reader can only relieve by reading on, and that is fine, because the theory the practice actually rests on is a different one: George Loewenstein's 1994 information-gap account of curiosity in Psychological Bulletin, which says that knowing a little about something makes you want the rest. Loewenstein was not tested by the 2025 paper and has not been overturned by it. Nobody should read the null on memory as a null on curiosity.
So the accurate charge is narrower than “your mechanism was debunked,” and better for being narrow. The industry borrowed a famous name for authority it did not need, from an effect about memory, to justify a practice about attention. And the effect that actually describes what it sells, that people come back to unfinished things, unforced, about two times in three, is the one it never learned to cite. The gamification writer Yu-kai Chou made the scientific half of this point in June 2026, reporting the 0.99 ratio and advising designers to bank on the resumption pull rather than the recall claim. The marketing half is still open, and it is the half with the money in it.
The industry also has better evidence available than a 1927 dissertation, and it is the industry's own. In 2025 Marianne Aubin Le Quéré and J. Nathan Matias published, in Scientific Reports, a meta-analysis of 8,977 headline experiments drawn from the Upworthy Research Archive, the time series of 32,487 A/B tests that the viral publisher ran on its own readers. The title is “When curiosity gaps backfire.” Headlines that are too vague lose clicks, as do headlines that are too concrete; the effect of a headline's concreteness depends on what the surrounding headlines look like. That is a real finding, from tens of thousands of real readers, about the actual lever. It says nothing about Bluma Zeigarnik, because it does not need to.
Why did the weaker finding become the famous one? Part of the answer is narrative. Zeigarnik's effect came with a story, the café waiter who could recite every unpaid order and forgot them the moment the bill was settled, and stories travel. Part of it is that a memory effect flatters the reader: it says your nagging thoughts about an unfinished project are your brain being clever, rather than your habit being ordinary. Resumption is behavioural and unglamorous. “People tend to go back and finish” does not sell a course.
And part of it is a small, concrete accident of bibliography. Maria Ovsiankina published her 1928 study under that name. By 1937 she was publishing in the Journal of General Psychology as Maria Rickers-Ovsiankina, on interrupted tasks in patients with schizophrenia, and her later career ran under the married name. A citation trail that changes surname mid-bibliography is a trail that breaks for anyone searching by the name they first met. It is a mundane reason for a finding to go uncited, which is exactly why it is worth naming: the record does not always pick winners by evidence.
If you build software, run a team, or write the copy that asks people to come back, the practical content of this paper fits in four moves.
First, retire the memory claim. If a deck, an onboarding doc or a growth plan says that unfinished tasks are remembered better, the current pooled estimate for that sentence is 0.99, and a reviewer who knows the 2025 paper will notice. If the argument needs “people are pulled back to unfinished work,” you have a better citation: the resumption finding, 67 percent, unforced, robust to dropping its founder.
Second, engineer resumability rather than withholding. The 67 percent is a return rate for tasks left in a recoverable half-done state, where the person can see what finishing would take. That is a design brief. The editor that reopens on the exact line you left. The pull request left with one named failing test as the next step. The checklist with three of five ticked. The draft that ends mid-sentence so tomorrow starts already in motion. None of these withholds anything from anyone. They lower the cost of coming back, which is the thing the surviving effect is about.
Third, notice what the inclusion rule excluded. Resumption counted only when the experimenter did not force it. A nagging notification is forced resumption, and the paper's own explanation for the vanished memory effect names phone notifications as the thing that destroyed task involvement. Every interruption you send to pull someone back spends down the state that would have made them care in the first place. If you want either effect, the lever is fewer interruptions, not more hooks.
Fourth, measure return, not recall. You almost certainly have no metric for how well users remember an unfinished flow, and it turns out there was no reason to want one. Return rates, resumed sessions, completed-after-abandoned funnels: the retention numbers you already have are measurements of the effect that replicated.
The Berlin laboratory produced two findings about unfinished work. The one everyone remembers is the one about memory, and it did not survive. The one that survived is about behaviour, and it explains something more useful than why an open loop nags. It explains why, two times in three, without anyone making you, you go back.
A citation trail that breaks mid-bibliography is a provenance failure. So is an agent's.
Ovsiankina's finding went uncited for decades partly because her name changed and the trail broke. The record did not pick winners by evidence; it picked them by what stayed traceable. That problem does not stay in the archives. When an autonomous agent produces a result, the question of which claim came from where, which step was checked, and which number is inherited rather than measured, is the same question with a shorter time horizon. Chain of Consciousness anchors each action to a tamper-evident record, so the provenance of a result is something you read off the chain instead of reconstructing from a trail that may already have broken.
See a verified action chain · Hosted Chain of Consciousness
pip install chain-of-consciousness · npm install chain-of-consciousness