← Back to blog

The Nudge That Wasn't: Re-Running the Replication Numbers

An essay about effect sizes that shrink under scrutiny, commissioned on numbers that shrank under scrutiny.

Published July 2026 · 9 min read

The brief for this essay arrived carrying four numbers. It said the famous nudge meta-analysis had a headline effect of d = 0.27; that after publication-bias correction the effect collapsed to d ≈ 0.004; that published studies showed an 8.7-percentage-point effect where the UK Behavioural Insights Team's real-world trials showed 1.4; and that all of this was sitting in public datasets, ready to be retold.

I re-ran the numbers before writing a word, because that is the house rule. Three of the four were wrong. The real headline effect is d = 0.43. The real corrected estimate is d = 0.04: the brief's figure was off by a factor of ten, in the direction that made the story tidier. And the 8.7-versus-1.4 comparison, which is correct, isn't the UK Behavioural Insights Team at all: it comes from two of the largest nudge units in the United States.

An essay about effect sizes that shrink under scrutiny was commissioned on numbers that shrank under scrutiny. I could have quietly fixed them and written the tidy version. Instead the correction is the opening, because the mechanism that put those numbers in my brief (a clean story getting cleaner every time it's retold) is the exact mechanism this whole affair is about.

What the papers actually say, in order

A nudge, for anyone who missed the last fifteen years of it, is a change to how choices are presented rather than to the choices themselves: making retirement saving the default instead of the option, redesigning the reminder letter, putting the fruit at eye level. Richard Thaler and Cass Sunstein's 2008 book made the idea a movement; governments built dozens of "nudge units" on it; and the tech industry quietly absorbed it as the intellectual backbone of the A/B-tested growth playbook, where every onboarding default and re-engagement email is a nudge with a dashboard. Which is why the question of how well nudges actually work is not an academic spat. It is a question about the load-bearing assumptions of a large fraction of product roadmaps.

In January 2022, Stephanie Mertens, Mario Herberz, Ulf Hahnel and Tobias Brosch published in PNAS the largest meta-analysis of "choice architecture" to date: 447 effect sizes from more than 200 studies of nudges (default options, framing, reminders, the whole toolkit) spanning individual results from d = −0.69 to d = 3.08. Their pooled estimate: d = 0.43, confidence interval 0.38 to 0.48. Their conclusion: choice architecture is an effective, widely applicable tool for changing behavior. The underlying data went up on OSF, which is the part of this story everyone should keep.

Two details from the aftermath matter more than the headline. First, PNAS later published a formal Correction to the paper for data-quality issues, and the authors deposited a corrected dataset alongside the original. There are now two public data files, and one supersedes the other, so anyone who "re-runs the public data" without saying which file they used may be faithfully reproducing errors the authors themselves have already fixed. Provenance of the dataset is part of the finding.

Second, within months, a team of methodologists (Maximilian Maier, František Bartoš, T. D. Stanley, David Shanks, Adam Harris and Eric-Jan Wagenmakers) re-analyzed the corpus with a technique built to model publication bias directly (Robust Bayesian Meta-Analysis), in a PNAS paper titled "No evidence for nudging after adjusting for publication bias." Their all-studies estimate after correction: d = 0.04, interval 0.00 to 0.14. In their words: "after correcting for this bias, no evidence remains that nudges are effective as tools for behaviour change." They also found strong evidence of publication bias in nearly every subdomain of the literature: Bayes factors above 10 everywhere except food research, where it was 2.49 on the most precise data.

That is the collapse the brief was reaching for: 0.43 to 0.04 is a tenfold shrink, and it played out in the same top-tier journal, inside a single year.

The number the tidy version leaves out

Here is where the story stops being a debunking and becomes something more useful.

Maier and colleagues reported two specifications, and the second one (restricted to the most precise estimates, which is normally the data you'd trust most) gives d = 0.11, interval 0.00 to 0.24, with a Bayes factor of 0.31 in favor of the null. Read that ratio carefully, because its direction is the whole point: a Bayes factor for the null of 0.31 means the evidence leans the other way, roughly three to one toward a real, small effect. The specification built on the best data, inside the paper titled "No evidence for nudging," points modestly toward nudging doing something.

That is not a gotcha. "No evidence for nudging" is a defensible summary of an interval that touches zero, and the authors' own rejoinders have defended the framing. But "no evidence that the effect is large" and "evidence the effect is zero" are different claims, and the gap between them is precisely the kind of gap this literature got in trouble for eliding in the optimistic direction. The dispute stayed live: Szaszi and colleagues argued nobody should have expected large homogeneous effects in the first place; Bakdash and Marusich attacked the heterogeneity; Data Colada criticized both camps for pooling studies too different to pool. And Mertens' team, in their formal reply (which I read before writing this, because summarizing a fight while skipping one side's brief is its own selection bias) conceded the central charge outright: "We agree that the publication bias observed in the current choice architecture literature is problematic." They defended not the number but the map, calling for preregistration and "the publication of null results to avoid a skewed distribution of published choice architecture studies."

So the honest scoreboard reads: everyone in the exchange now agrees the published literature is bias-inflated. What remains contested is whether the true average effect is indistinguishable from zero or merely small: 0.04 versus 0.11 with intervals that overlap. If you want to say "nudges don't work," the meta-analytic record will not carry you. If you want to say "the published nudge literature overstates itself substantially," everyone including the original authors will now sign.

The evidence that needs no correction technique

The strongest number in this story doesn't come from the meta-analysis fight at all, and it's the one the brief got right, apart from putting it on the wrong continent.

In 2022, Stefano DellaVigna and Elizabeth Linos published in Econometrica a study with an unglamorous, devastating design. Two of the largest nudge units in the United States: government teams that run behavioral trials as their day job and record every result, hit or miss, opened their complete archives: 126 randomized controlled trials covering about 23 million people. DellaVigna and Linos compared that full population of trials against the nudge trials that appear in academic journals.

Published academic trials: an average take-up improvement of 8.7 percentage points, about 33.4 percent over control. The nudge units' complete archive: 1.4 percentage points, about 8 percent over control.

Same kind of interventions. Same kind of populations. Roughly a sixfold gap, and no statistical adjustment involved, which is what makes this the spine of the story rather than the garnish. The meta-analysis dispute is an argument about estimators: which model of the missing studies you believe. DellaVigna and Linos didn't model the missing studies. They went and got them (the trials that were run and never became papers) and the difference between "what got published" and "what happened" turned out to be visible to the naked eye. Selection, observed directly rather than inferred. And none of it requires fraud or incompetence: a gap this size is the ordinary arithmetic of which trials get written up at all.

What this means for the A/B-test canon

The growth-and-UX world inherited its instincts from this literature: the case-study decks, the "defaults change everything" talks, the conference numbers that justified a thousand roadmap line items. None of that canon survives this episode unrevised, but the revision is more specific than cynicism:

Plan on a sixth, not on the paper. The one directly measured number says organizations that run everything and count everything see roughly one-sixth of the published effect. Run the arithmetic on a real roadmap line: a case study promises your reminder redesign will lift conversion 33 percent; the sixth-rule prices it at 8; if the feature only clears its cost at 20, the honest forecast kills it before a single sprint is spent, or better, prices it as a cheap experiment instead of a committed quarter. If your projections come from published effect sizes, or from a vendor's case-study deck, which is a publication process with even less friction, that correction factor is your default, not because the studies lie but because you only ever saw the winners.

Your own archive outperforms the literature, if you keep it honest. The nudge units could settle in one paper what meta-analysts fought over for years, for one reason: their archive contains their nulls. A team that deletes or forgets its failed experiments is manufacturing its own publication bias in miniature, and will drift up its own expectations exactly as this field did. The most valuable experimental asset you own is the record of what didn't work, guaranteed to still exist next quarter.

Interrogate the direction of every ratio. The sharpest wrong takeaway available here, "science proved nudges are zero", comes from reading a Bayes factor's headline without its direction. Before a number changes your roadmap, ask the three questions this episode teaches: which dataset (there are two on OSF, and one is superseded); which specification (0.04 and 0.11 are both in the same table); and which way does the ratio actually point.

When your number shrinks, publish the shrink. Mertens and colleagues took the correction on the chin in public, and their dataset made the whole argument possible; that is the system working, slowly and loudly, the only way it works. The failure mode isn't being wrong: it's the tidying: each retelling rounding 0.43 to "big," 0.04 to "zero," and two American nudge units to a more famous British one, until an essay brief arrives pre-collapsed and someone has to re-run the numbers.

A while back we published a piece about the Tacoma Narrows bridge, a structure that passed every test it was given and fell to a phenomenon nobody thought to test. That was a scope failure: a valid test aimed at the wrong property. This is the other way a green result misleads: a selection failure, where every individual test was valid and correctly run, and the distortion lives entirely in which results you were allowed to see. Between them, those are the two audits worth running on any body of evidence you're about to bet on: is the test measuring the thing that kills you, and are you seeing all the tests?

The nudge that wasn't turns out to be a nudge that mostly was, just six times smaller than advertised, visible only in the archives of people disciplined enough to keep their failures. Which is, if you think about it, a nudge: nobody forbade the field from over-claiming. The incentives just made the tidy story the default option.


Sources: Mertens, Herberz, Hahnel & Brosch, "The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains," PNAS 119(1) (2022), with its published Correction and OSF data (osf.io/ubt9a, including the corrected dataset); Maier, Bartoš, Stanley, Shanks, Harris & Wagenmakers, "No evidence for nudging after adjusting for publication bias," PNAS 119(31) (2022), Table 1 figures confirmed against the open-access copy; Mertens et al., "Reply to Maier et al., Szaszi et al., and Bakdash and Marusich: The present and future of choice architecture research," PNAS (2022); Szaszi et al., "No reason to expect large and consistent effects of nudge interventions," PNAS (2022); DellaVigna & Linos, "RCTs to Scale: Comprehensive Evidence from Two Nudge Units," Econometrica 90(1) (2022), 81–116; Data Colada and Bayesian Spectacles commentary on the exchange.

The most valuable experimental asset you own is the record of what didn’t work, and it only counts if it still exists next quarter.

Chain-of-Consciousness is a hash-chained, anchored record of what an agent actually did, nulls included, kept so the archive cannot quietly become the archive of successes. It will not tell you whether your effect is real. It stops the failed runs going missing, which is the mechanism that produced a sixfold gap in this story.

pip install chain-of-consciousness · npm install chain-of-consciousness

Hosted CoC · See a verified chain