TL;DR: Across 126 randomized trials covering roughly 23 million participants, the two largest US government nudge units measured an average effect of 1.4 percentage points on a base rate near 17 percent, about 8 percent relative. The comparable academic literature reports 8.7 percentage points, roughly six times larger. DellaVigna and Linos (2022, Econometrica) attribute a large share of that gap to publication bias and most of the remainder to trial characteristics, above all sample size. The operating conclusion is not that nudges fail. It is that the effect size a growth team should budget for is roughly an order of magnitude smaller than the trade literature implies, which changes both which interventions are worth building and how large a test has to be before its result means anything.
The Lift That Never Arrives
The pattern is common enough that I now recognize it from the calendar invite. A growth or lifecycle team has read a behavioral science book, or attended a conference talk, or hired someone who did. They have identified a place in the funnel where a well-designed nudge should work: a social-proof line on the checkout page, a loss-framed reminder email, a default pre-selection on the shipping options. They build it. They ship it to a randomized holdout. Two weeks later the readout comes back, and the treatment effect is a rounding error with a confidence interval wide enough to contain both a meaningful win and a meaningful loss. Somebody says the test was inconclusive. Somebody else says the implementation was wrong. The intervention gets shipped anyway, or gets killed, and the decision is made on the basis of whoever spoke last.
The prior and the test fail together
What happened is not mysterious, and it is not usually an implementation failure. Two things went wrong at once, and they compound. The first is that the team's prior on the effect size was calibrated against a literature that systematically over-reports. The second is that the test was sized to detect the over-reported effect, which means it had almost no chance of detecting the real one. A team expecting a five-point lift will size a test for a five-point lift. If the true effect is one and a half points, that test will return "no significant difference" with high probability, and the organization will learn nothing except that behavioral science is overrated.
Both failures trace to the same root cause, which is that the published nudge literature and the operating reality of nudges at scale are describing different distributions. The academic literature contains the trials that produced publishable results. The operating world contains all the trials. Those are different samples, and until 2022 there was no clean way to see how different.
There is now. The purpose of this essay is to work through what the corrected estimate actually is, why the correction is as large as it is, what survives the correction, and how a team should change its experiment design and its budgeting as a result. The short version of the answer is that the correct expected effect for a messaging-class behavioral intervention is on the order of one to two percentage points on a base rate in the teens, that most tests run by most growth teams are badly underpowered to detect that, and that the appropriate response is to build more measurement capacity rather than to abandon the intervention class.
The Cleanest Benchmark We Have
The single most useful paper in this literature is DellaVigna and Linos (2022), "RCTs to Scale: Comprehensive Evidence from Two Nudge Units," published in Econometrica 90(1), 81–116, doi:10.3982/ECTA18709. Its design solves the selection problem directly rather than trying to model around it.
The authors obtained the complete trial portfolios of the two largest nudge units operating in the United States: the North American arm of the Behavioural Insights Team (bi.team) and the federal Office of Evaluation Sciences (oes.gsa.gov). Between them, these organizations had run 126 randomized controlled trials covering roughly 23 million participants across a range of government programs: benefit take-up, tax compliance, vaccination, college enrollment, health-service scheduling, and similar. The average trial therefore involves something on the order of 180,000 people, arithmetic on the reported totals rather than a figure the authors emphasize, but it conveys the scale difference correctly.
The headline result: the average treatment effect across those 126 trials was approximately 1.4 percentage points, against an average control-group take-up rate of roughly 17 percent. That is about an 8 percent relative improvement. It is a real effect, statistically distinguishable from zero when pooled, and in many government contexts it is cost-effective, because a text message costs almost nothing and 1.4 points of additional benefit take-up across a population of millions is a substantial number of people.
Then the comparison. DellaVigna and Linos benchmarked this against the academic nudge literature, using as the comparison set the trials collected in Benartzi, Beshears, Milkman, Sunstein, Thaler, Shankar, Tucker-Ray, Congdon and Galing (2017), "Should Governments Invest More in Nudging?", Psychological Science 28(8) — a review assembled specifically to make the case that nudges are a high-return policy instrument. The average effect in that published academic set was approximately 8.7 percentage points. That is roughly 6.2 times the nudge-unit average.
Why the Nudge Units Are the Right Benchmark
The comparison only carries weight because of an institutional detail that is easy to skim past. The nudge units report every trial they run.
This is not a courtesy. It is a structural feature of how these organizations are set up. The Office of Evaluation Sciences publishes its evaluations regardless of outcome, including the ones where the intervention did nothing, because its mandate is to inform federal agencies about what works rather than to generate publications. The Behavioural Insights Team's North American portfolio was made available to the researchers in full. There is no file drawer, because there is no incentive to maintain one. A null result in a nudge unit is a finding to be reported to the client agency, not a failure to be quietly shelved.
That single property is what makes the 1.4-point figure a fundamentally different kind of estimate than the 8.7-point figure. One is an average over all trials attempted. The other is an average over the trials that cleared an editorial and career filter which strongly prefers statistically significant, large, novel results. Comparing them is not comparing two estimates of the same quantity. It is comparing a population mean to a mean conditional on selection.
DellaVigna and Linos do not stop at the raw comparison. They decompose the gap. Their analysis attributes a large share of it to publication bias in the academic literature, which they model explicitly using a selective-publication framework, and most of the remainder to observable differences in trial characteristics between the two sets. The dominant trial characteristic is sample size. Nudge-unit trials are enormously larger than academic trials, and within their data, larger trials produce smaller estimated effects. That relationship is not incidental. It is the second half of the mechanism, and it deserves its own treatment.
One further characteristic matters for how a practitioner should read the 1.4-point number, and it cuts in the direction of caution about generalizing it. The nudge-unit portfolio skews heavily toward low-cost messaging interventions: letters, emails, text messages, reminders, simplified forms, social-comparison framing. These are the interventions a government agency can deploy at scale without legislative change or systems rework. Structural interventions — changing a default, changing what the form requires, changing the order of operations in an enrollment flow — are comparatively rare in the portfolio, because they are comparatively hard to authorize. The 1.4-point average is therefore best read as the expected effect of a messaging-class nudge, not as the expected effect of any behavioral intervention whatsoever.
Two Evidence Bases on Nudge Effectiveness, Compared
| Property | Two US nudge units (DellaVigna and Linos 2022) | Published academic nudge RCTs (Benartzi et al. 2017 set) |
|---|---|---|
| Number of trials | 126 | Smaller set, assembled from published journal articles |
| Total participants | Approximately 23 million | Orders of magnitude smaller in aggregate |
| Average participants per trial | Roughly 180,000 (arithmetic on reported totals) | Typically hundreds to low thousands |
| Average measured effect | About 1.4 percentage points | About 8.7 percentage points |
| Average control take-up | About 17 percent | Varies by domain |
| Relative effect | About 8 percent | Substantially larger |
| Selection into the sample | Census: all trials reported regardless of outcome | Conditional on publication, which favours significant and large results |
| Dominant intervention type | Messaging: letters, emails, texts, reminders, simplification | Mixed, with more structural and laboratory-adjacent designs |
| What the average estimates | Population mean of attempted interventions | Mean conditional on passing the significance filter |
Why Larger Trials Find Smaller Effects
The inverse relationship between sample size and estimated effect is one of the most reliable regularities in applied social science, and it is also one of the most frequently misread. Practitioners tend to interpret it as evidence that something goes wrong when interventions scale — that the magic wears off, that the intervention gets diluted through a larger and less receptive population. That interpretation is partly right and mostly wrong. Three distinct mechanisms produce the pattern, and they have very different implications for what an operator should do.
Mechanism One: The Significance Filter
The first mechanism is purely statistical and it is the largest. It has nothing to do with the intervention and everything to do with what gets reported.
The mechanism works like this. Suppose the true effect of some intervention is 1.4 percentage points on a 17 percent base rate. A researcher runs it on 500 people per arm. The standard error of the estimated difference in that design is about 2.4 percentage points. For the result to reach conventional significance at the five percent level, the observed effect has to be at least roughly 1.96 standard errors from zero, which is about 4.7 percentage points. The true effect is 1.4. Therefore the only way this study produces a publishable result is if random sampling variation happens to push the estimate more than three times above the truth.
It follows mechanically that every significant estimate from that design is a large overstatement. There is no exception, and it is not that low-powered studies are noisy in both directions and the average comes out right: conditional on significance they are biased upward with certainty, because the threshold is a floor on the reported magnitude. This is the winner's curse applied to research: the estimates that win the publication auction are the ones that overbid.
The magnitude of the distortion can be calculated exactly. Assume a true effect of 1.4 percentage points on a 17 percent base rate, a two-arm equal-allocation design, and a two-sided test at the five percent level. Then the following holds.
Statistical Power and the Exaggeration Ratio for a True 1.4-Point Effect on a 17 Percent Base Rate
| Sample per arm | Standard error | Power to detect the true effect | Expected reported effect if significant | Exaggeration ratio | Probability the sign is wrong |
|---|---|---|---|---|---|
| 500 | 2.38 pp | 9 percent | 5.7 pp | 4.1x | 6.0 percent |
| 1,000 | 1.68 pp | 13 percent | 4.1 pp | 3.0x | 2.0 percent |
| 2,000 | 1.19 pp | 22 percent | 3.0 pp | 2.2x | 0.4 percent |
| 5,000 | 0.75 pp | 46 percent | 2.1 pp | 1.5x | Under 0.1 percent |
| 10,000 | 0.53 pp | 75 percent | 1.6 pp | 1.2x | Negligible |
| 20,000 | 0.38 pp | 96 percent | 1.4 pp | 1.0x | Negligible |
| 50,000 | 0.24 pp | Over 99 percent | 1.4 pp | 1.0x | Negligible |
These figures are computed directly from the normal approximation to the two-proportion test, conditioning the estimate on crossing the two-sided five percent threshold; they are arithmetic, not empirical findings, and they hold whatever the substantive domain. The column that matters is the fifth one. A study with 500 per arm that reports a significant result on a true 1.4-point effect will, on average, report something above five and a half points — which is very close to the 8.7-point academic average once you allow for the fact that many published trials are smaller still and that some domains have larger true effects. The published literature is not reporting a different effect. It is reporting the same effect through a filter that multiplies it by three or four.
The sixth column deserves a mention too, because it is the part practitioners almost never think about. At 500 per arm, roughly six percent of the significant results will have the wrong sign — they will report a negative effect for an intervention that genuinely helps, or vice versa. Gelman and Carlin call this a Type S error. In a small-sample experimentation program, a non-trivial fraction of the "we tested it and it hurt conversion" conclusions are noise pointing the wrong way.
This is not a hypothetical concern about the social sciences in general. Franco, Malhotra and Simonovits (2014, Science) examined the Time-sharing Experiments for the Social Sciences archive, which records every study fielded through the program regardless of what it found, and demonstrated that strong results were dramatically more likely to be written up and published than null results — with a large majority of the null findings never even drafted into a paper. The Open Science Collaboration (2015, Science) attempted direct replication of 100 psychology studies and recovered significant effects in around a third, with replication effect sizes roughly half the originals. Camerer and colleagues (2018, Nature Human Behaviour) replicated social science experiments published in Nature and Science and found a similar pattern: most replicated in direction, but at roughly half the original magnitude. Halving is what you get when the original literature is moderately powered. Sixfold shrinkage is what you get when it is badly powered and the pooled comparison set is a policy-advocacy review rather than a random sample.
Mechanism Two: Genuine Heterogeneity and Site Selection
The second mechanism is real and substantive: effects genuinely differ across contexts, and the contexts chosen for early, small, academic trials are systematically the ones where the effect is largest.
The cleanest demonstration of this outside the nudge literature proper comes from Allcott (2015), "Site Selection Bias in Program Evaluation," Quarterly Journal of Economics. Allcott studied the Opower home energy report program, which mails households a comparison of their electricity use against similar neighbours — a canonical social-comparison nudge, and one of the most heavily evaluated behavioral interventions in existence. The 2015 paper showed that the effect measured in the program's first sites was materially larger than the effect measured as it expanded to a broader and more representative set of utilities. The early sites were not representative. They were selected — by the vendor, by the utilities that volunteered first — in ways correlated with responsiveness.
This is the honest version of "it does not scale," and it has a specific structure. The pilot is run where conditions are favourable: an engaged population, a motivated operator, a well-implemented intervention, a context where the friction being removed was actually binding. Scale means running the same intervention where those conditions do not hold. The average across the expanded population is lower, and it is lower for real reasons rather than statistical ones.
Szaszi, Higney, Charlton, Gelman, Ziano, Aczel, Goldstein, Yeager and Tipton (2022), "No reason to expect large and consistent effects of nudge interventions," PNAS 119(31), makes the general form of this argument forcefully. Their position is that "nudge" is not a treatment with a stable effect size, because it is not a treatment at all — it is a loose family of interventions deployed across domains and populations with wildly different baseline motivations. The expected effect of a nudge, averaged over that space, is not a scientifically meaningful parameter. It is a number produced by an aggregation choice. Asking "how big is a nudge" is, on this view, like asking how big a medical intervention is.
I think this critique is correct and I also think it is more useful to practitioners than it first appears. It does not say the evidence is unusable. It says the unit of analysis has to be the specific mechanism in the specific context, and that the pooled average is a prior to be updated rather than a forecast.
Mechanism Three: Implementation Dilution
The third mechanism is the most mundane and the one operators most reliably underestimate. At small scale, the intervention is delivered by the people who designed it. At large scale, it is delivered by a system. A researcher-designed reminder letter is proofread, tested for comprehension, timed carefully, and sent to a clean list. The same letter at scale goes out through an existing mail merge with a stale address file, arrives alongside three other communications from the same agency, and lands in a week when the call centre is understaffed. None of this appears in the treatment description; all of it reduces the measured effect. It is one reason the nudge-unit trials, which run inside real operational systems, produce estimates closer to what a practitioner will actually experience than a research team's own pilot does.
The Mertens and Maier Exchange
If the DellaVigna and Linos paper establishes the size of the correction, a two-step exchange in PNAS during 2022 establishes how completely a publication-bias correction can dissolve a headline finding. It is the sharpest single illustration available of the problem, and it is worth walking through in detail because the methodological point generalizes well beyond nudging.
In January 2022, Mertens, Herberz, Hahnel and Brosch published "The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains" in PNAS 119(1). It was a large, competent, conventionally executed meta-analysis of the choice-architecture literature, and it reported a positive overall effect of approximately Cohen's d = 0.43. It was widely reported as a vindication: nudging works, the meta-analysis confirms it, here is the number.
In August 2022, in the same journal, Maier, Bartoš, Stanley, Shanks, Harris and Wagenmakers published "No evidence for nudging after adjusting for publication bias," PNAS 119(31), doi:10.1073/pnas.2200300119. They did not collect new data. They reanalyzed the same dataset using Robust Bayesian Meta-Analysis, a method that does not assume the absence of publication bias but instead models the selection process explicitly, averaging over a family of models that includes both bias-present and bias-absent specifications and weighting them by how well each explains the observed distribution of results.
The result was not a modest downward revision. After adjustment, the evidence for an overall nudging effect essentially disappeared. The corrected pooled estimate fell to approximately zero, and — this is the part that distinguishes the paper from a routine caveat — the Bayes factor indicated evidence for the absence of an effect rather than merely an absence of evidence for one. The distinction is not pedantic. A wide confidence interval containing zero means the study could not tell. A Bayes factor favouring the null means the data actively support the hypothesis that the pooled effect is nil, once the selection process is accounted for.
Szaszi and colleagues, in the paper cited above, published in the same issue and made the complementary argument: that even setting aside publication bias, the heterogeneity across nudge interventions and contexts is severe enough that a single pooled effect size is the wrong summary statistic. Mertens and colleagues replied in the same venue. The exchange did not converge on an agreed number, and I do not think it should have, because the disagreement is partly about whether the pooled number is a coherent target at all.
Two things follow for a practitioner, and they point in slightly different directions.
The first is straightforwardly deflationary. If a headline meta-analytic effect can fall from d = 0.43 to approximately zero purely by changing the model of publication selection, then the headline effect in any behavioral literature you have not personally checked for bias correction should be treated as an upper bound rather than an estimate. Most of what appears in practitioner books, conference decks, and vendor case studies falls under this rule, and none of it runs a bias correction.
The second is a caution against over-reading in the other direction. "No evidence for an overall nudging effect after correction" is not the same claim as "no nudge works." It is a claim about the pooled average across a heterogeneous category — precisely the quantity Szaszi and colleagues argue is not meaningful. The nudge-unit census gives us something the meta-analysis cannot: an unfiltered average of 1.4 percentage points that is not zero, measured on real interventions in real programs at real scale. The reconciliation is that the truth about the messaging-class nudge is small, positive, and highly variable — which reads as "approximately zero" when you pool across an incoherent category with heavy selection, and as "1.4 points" when you take a census of a specific one.
A headline effect that survives only under the assumption of no publication bias is not a finding. It is a hypothesis about the file drawer.
Megastudies as the Methodological Fix
The methodologically strongest response to all of this is not better meta-analysis. It is a different experimental design, and it already exists.
A megastudy tests many interventions simultaneously, in one population, under one preregistered protocol, against a common control, and reports every arm. The structure removes the file drawer by construction: there is no decision point at which a non-performing arm can be dropped, because all arms were registered together and the paper reports the full ranking. It also removes the cross-study comparability problem, since every intervention faces the same population, the same outcome measure, the same time window, and the same delivery infrastructure. The remaining differences between arms are attributable to the interventions themselves rather than to the settings in which they were tested.
Most arms do almost nothing
Milkman and a large collaborating group published the first at scale in Nature 600 (2021), "Megastudies improve the impact of applied behavioural science," which tested 54 different interventions on 61,293 members of a national gym chain over a four-week program. The results are instructive in three ways. Most interventions produced small effects. The distribution of effects across arms was tight enough that the difference between a good intervention and a mediocre one was much smaller than most practitioners would guess. And — this is the finding least often quoted — the effects largely faded once the intervention period ended. Behaviour during the treatment window moved; behaviour after it mostly reverted.
The vaccination megastudies followed a similar shape. Milkman and colleagues (2021, PNAS) tested text-message nudges encouraging patients to get a flu shot at an upcoming doctor's appointment; the best-performing arm raised vaccination rates by roughly 11 percent relative to the holdout, and most arms did considerably less. A subsequent, much larger pharmacy megastudy (Milkman and colleagues, 2022, PNAS) tested 22 text-message interventions across 689,693 pharmacy patients; the best-performing message produced a single-digit relative improvement, and the average across all 22 arms was far smaller than that.
The design also exposes something the published literature could never show: how poorly expert intuition predicts which arm wins. Where forecasts were collected alongside the trials, predicted rankings tracked the realised rankings only loosely. That is not a criticism of the forecasters. It is a direct consequence of a distribution in which most true effects are small and clustered — when the real differences between interventions are on the order of a percentage point, no amount of theoretical sophistication will let you rank them by inspection. You have to measure.
For an operator, the megastudy is both a benchmark and a template. As a benchmark, it says: when 22 or 54 competent interventions are tested honestly in one population, the results look like this. As a template, it says: if you have the traffic, testing many variants against one control in a single preregistered wave is a better use of your experimentation budget than testing one variant per quarter, because it produces a distribution rather than a point, and the distribution is what you need for budgeting.
What Actually Survives
The correction is not uniform across intervention types, and this is where the practical value of the literature sits. Some behavioral mechanisms hold up considerably better than the pooled average, and there is a structural reason why.
Defaults are the clearest case. Madrian and Shea (2001, Quarterly Journal of Economics) documented what happened when a large employer switched its 401(k) plan from opt-in to automatic enrollment: participation among newly hired employees rose from roughly 37 percent to roughly 86 percent. That is not a 1.4-point effect. It is a structural change in what the employee has to do to end up in the plan, and the magnitude has held up across many subsequent implementations, which is not something one can say about most famous behavioral findings.
The mechanism is different in kind from a messaging nudge, and the difference explains the persistence. A social-proof line on a checkout page changes the salience of information; the customer still has to take exactly the same action to convert. A default changes which outcome occurs when the person takes no action at all. Since a large fraction of people take no action at all in most flows, moving the no-action outcome moves a large fraction of the population directly. There is no persuasion step to fail.
What the defaults evidence does not support
Three caveats are necessary, and I want to be careful here because "defaults work" has itself become a slogan that outruns its evidence.
The first is that a default is only as powerful as the passivity it exploits. Where the decision is high-stakes, well-understood, and actively deliberated, defaults do much less. A default that reallocates a customer's entire budget will be noticed and overridden; a default that pre-selects a shipping speed will not.
The second is that registration is not outcome. The organ donation literature is the standard illustration and also the standard warning: the very large differences in registered or presumed consent between opt-in and opt-out countries do not translate proportionally into differences in actual transplants performed, because the binding constraints downstream — family consent conversations, clinical eligibility, organ procurement logistics — are unaffected by the default. The default moved the metric it directly touched. It moved the outcome the metric was standing in for much less: jurisdictions that have switched to deemed consent have generally not seen increases in donation rates commensurate with the change in registered consent.
The third is that the nudge-unit portfolio contains relatively few default changes, which means the 1.4-point census average is not a good estimate of the default effect specifically — it is a good estimate of the messaging effect. We do not have a comparably clean census for defaults, and I would treat any confident numerical claim about the average default effect with the same skepticism this essay has directed at the average nudge effect.
Between the two poles sit mandated-choice and active-decision mechanics, which force a person to make an explicit selection before proceeding rather than defaulting them either way. Carroll, Choi, Laibson, Madrian and Metrick (2009, Quarterly Journal of Economics, "Optimal Defaults and Active Decisions") showed that requiring new hires to make an explicit enrollment decision raised participation substantially relative to standard opt-in, without the paternalism objection that automatic enrollment attracts. Mechanically, these interventions work for the same reason defaults do: they change the structure of the choice rather than the framing of the information.
Interventions That Change the Structure of the Choice, Ranked by How Well They Survive Bias Correction and Scale
| Family | What it changes | Evidence status | Practitioner prior to budget |
|---|---|---|---|
| Defaults and automatic enrollment | The outcome that occurs when the person does nothing | Strong. Large effects replicated across many implementations. Weakest where decisions are high-stakes and actively deliberated. | Large, but only in flows with high passivity; verify the metric is the outcome and not a registration proxy |
| Mandated choice and active decision | Forces an explicit selection before proceeding | Good. Smaller literature than defaults but consistent direction. | Moderate; also usually easier to authorize than a default change |
| Friction removal and simplification | The number of steps and the cognitive cost of completing them | Reasonable. Well represented in the nudge unit portfolio and among its better performers. | Small to moderate; scales with how much friction actually exists |
Interventions That Change the Information or Its Framing, Ranked the Same Way
| Family | What it changes | Evidence status | Practitioner prior to budget |
|---|---|---|---|
| Reminders and timing | When the prompt arrives relative to the decision window | Reasonable. The workhorse of the nudge unit portfolio. | Small, roughly the census average; cheap enough that small still pays |
| Planning prompts and implementation intentions | Whether the person has articulated a concrete plan | Mixed. Positive in several megastudy arms, unimpressive in others. | Small and variable |
| Social proof and social comparison messaging | The salience of what other people do | Weak to mixed after correction. Effects present at scale in energy use, but subject to documented site selection bias. | Small; do not budget the published figures |
| Loss framing and gain framing | The wording of an identical offer | Weak after correction. Heavily represented in the small-sample literature most affected by the significance filter. | Near zero as a central estimate; treat any measured win as provisional |
The right-hand column of those tables is a practitioner prior rather than a measurement, reflecting how often each family shows up as a meaningful driver in advisory partner data and best read as an ordering with rough magnitudes rather than calibrated effect sizes; the evidence-status column is anchored to the published literature cited throughout this essay. I have deliberately not attached numbers to the prior, because attaching numbers would reproduce the exact error the essay is about.
The Case for More Infrastructure, Not Less Nudging
The distinction between knowing and not knowing is worth more than it sounds, because the economics of a messaging-class intervention are asymmetric in a way the economics of a product feature are not. A reminder email costs almost nothing to send and almost nothing to maintain. If it produces a 1.4-point lift on a 17-point base, that is an 8 percent relative improvement on a line item, achieved at effectively zero marginal cost, and it compounds with every other such improvement in the funnel. The problem was never that the effects are too small to be worth having. The problem is that they are too small to be seen by the measurement apparatus most teams have, which means the team cannot distinguish the ten that work from the forty that do not, and ends up maintaining all fifty or none. Ten independent 8 percent relative improvements are not an 80 percent improvement, but they are a material one.
So the correct response to the effect-size collapse is to invest in the thing that makes small effects legible: larger test populations per arm, longer runtimes, pooled tests that share one control across many treatments in the megastudy pattern, holdout groups maintained over quarters rather than weeks, and a standing prior that says a messaging intervention should be assumed to produce one to two points until proven otherwise. Teams that cut their experimentation programme in response to disappointing results are optimizing in exactly the wrong direction. They are removing the instrument rather than recalibrating it.
There is a second-order argument as well. An organization that does not measure small effects reliably does not simply lose information; it fills the gap with narrative. Interventions get shipped because a senior person likes them, kept because nobody has evidence against them, and cited in strategy documents as proven wins. The accumulated cost of that is not the forgone lift on any single test. It is a roadmap increasingly populated by things that were never actually validated, and an organizational memory that has no error-correction mechanism.
A Budgeting Framework
The practical output of everything above is a way of deciding, before building, whether a proposed behavioral intervention is worth the engineering time and whether the resulting test can possibly answer the question. Four steps.
Step One: Set the Prior from the Census, Not the Literature
Start from 1.4 percentage points on the relevant base rate for anything in the messaging family — reminders, framing, social proof, subject-line changes, on-page copy. That is the DellaVigna and Linos census average, and it is the least selected number available. Adjust up if the intervention is structural (a default, a mandated choice, a genuine friction removal), adjust down if the base rate is already high or the population has seen similar messaging repeatedly. Do not adjust up because a published paper or a vendor case study reports a large effect; that is precisely the number the correction applies to.
Express the prior as a range rather than a point, and make the range wide. Genuine heterogeneity is large, so a defensible prior for a messaging intervention might span roughly zero to three points with a central estimate near one and a half. The width is the honest part.
Step Two: Check Whether the Test Can See the Prior
This is the step teams skip, and skipping it is what generates the inconclusive readouts. The minimum detectable effect at 80 percent power, on a base rate of 17 percent, with a two-sided five percent test and equal allocation, is as follows.
Read from the other direction, the sample requirements are these.
Sample Size Required per Arm to Detect a Given Absolute Lift on a 17 Percent Base Rate at 80 Percent Power
| Absolute lift to detect | Relative lift | Sample per arm | Total sample | Comment |
|---|---|---|---|---|
| 0.5 pp | 3 percent | About 89,600 | About 179,000 | Realistic for a weak messaging intervention; out of reach for most teams |
| 1.0 pp | 6 percent | About 22,700 | About 45,300 | The threshold a serious lifecycle programme should be able to hit |
| 1.4 pp | 8 percent | About 11,700 | About 23,300 | The nudge unit census average |
| 2.0 pp | 12 percent | About 5,800 | About 11,600 | Typical upper end of what a mid-sized operator can run |
| 3.0 pp | 18 percent | About 2,600 | About 5,300 | Requires an unusually strong intervention to be a real target |
| 5.0 pp | 29 percent | About 980 | About 2,000 | Structural change territory, not messaging |
| 8.7 pp | 51 percent | About 344 | About 690 | The published academic average; a test this size is a coin flip dressed as evidence |
The bottom row is the whole problem in one line. A test with 344 people per arm is adequately powered if and only if the published academic effect is the true effect. It is the sample size a team will naturally choose if it has calibrated its expectations on the trade literature. And it is the sample size at which, per the exaggeration table above, any significant result it produces will overstate the truth by a factor of four and will have a meaningful chance of pointing the wrong way.
Step Three: Change What You Build, Not Just How You Test It
The most important consequence of a corrected prior is that it re-ranks the backlog. Under a 1.4-point prior for messaging and a structurally larger prior for defaults, a copy change and a checkout default are not comparable bets at all. At 8.7 points the two look like comparable bets and the copy change wins on engineering cost; under the corrected prior the calculation flips, and it is worth spending five times the engineering effort on the intervention that changes the no-action outcome, because the expected effect is several times larger and far more durable.
Concretely, this pushes a backlog toward: default selections and pre-population wherever the decision is genuinely low-stakes; mandated choice where a default would be presumptuous or regulated; removal of steps, fields and confirmation screens; and timing changes that place a prompt inside the window when the decision is actually being made. It pushes away from: framing variants of an identical offer, social-proof copy without an underlying behavioural claim, and the long tail of subject-line and button-colour tests that consume experimentation capacity while sitting entirely below the detection threshold.
That last point is worth stating plainly, because it is the discipline most teams find hardest. Tests that cannot detect their own realistic effect should not be run at all. They are not free — they consume traffic and analyst attention, they occupy a slot in the testing queue, and they generate spurious winners that get shipped and then get cited. A programme that runs twelve adequately powered tests a year learns more than one that runs sixty underpowered ones, and it learns things that stay true.
Step Four: Pool Your Tests
If your traffic will not support an adequately powered test of one intervention, it will certainly not support sixty. The megastudy structure is the response: share one control group across many treatment arms, register the full set in advance, run them concurrently rather than sequentially, and report every arm.
This buys three things. It removes the sequential-testing problem, in which a team runs variants one after another until something crosses significance — a significance filter operating inside your own organization. It makes the control efficient, since one large control serves all arms rather than each test allocating its own. And it produces a distribution rather than a point estimate, which is the input the budgeting step actually needs: after one wave you know not just whether an intervention worked but where it sits in the range of things you tried. The multiple-comparisons cost is real and must be handled, but it is a known problem with known solutions, and a considerably smaller one than it replaces.
An underpowered experiment does not produce a weak answer. It produces a strong answer that is wrong by a factor of three.
What This Does Not Say
Three misreadings of this argument are common enough to be worth pre-empting.
It does not say nudges do not work. The census evidence — 126 trials, 23 million people, every trial reported — shows a positive average effect that is statistically distinguishable from zero and frequently cost-effective. The Maier and colleagues result is a statement about a pooled average over a heterogeneous category with heavy selection, and the correct reading of it alongside DellaVigna and Linos is that the true distribution of nudge effects is centred on a small positive number with wide dispersion, not that it is centred on zero.
It does not say published research is untrustworthy in general. It says that effect magnitudes drawn from small-sample literatures are systematically inflated, by an amount that can be calculated from the sample sizes involved, and that the inflation is a property of the reporting filter rather than of any individual study.
And it does not say that behavioral science is the wrong lens. The mechanisms are real. Defaults exploit passivity, friction suppresses completion, timing determines whether a prompt lands inside a decision window. What the effect-size literature corrects is not the mechanism but the magnitude, and magnitude is the input to every resourcing decision a team makes. Getting the mechanism right and the magnitude wrong by a factor of six produces confident, well-reasoned, systematically bad prioritization.
Key Takeaways
-
The best available unfiltered estimate of a messaging-class nudge effect is approximately 1.4 percentage points on a base rate near 17 percent — about 8 percent relative — measured across 126 trials and roughly 23 million participants by two US nudge units that report every trial they run (DellaVigna and Linos, 2022, Econometrica). The comparable published academic literature reports about 8.7 percentage points, roughly six times larger, and the gap is attributed largely to publication bias with most of the remainder tracking trial characteristics, principally sample size.
-
Larger trials find smaller effects for three separable reasons, and only one of them is about the intervention. The dominant reason is the significance filter: a low-powered study can only reach significance when noise pushes the estimate far above the truth, so every published estimate from such a study is inflated by construction. Genuine context heterogeneity and site-selection bias (Allcott, 2015, QJE) and implementation dilution at scale account for the rest.
-
The exaggeration arithmetic is computable and stark. For a true effect of 1.4 points on a 17 percent base rate, a test with 500 per arm has about 9 percent power, will report roughly 5.7 points when it reaches significance — an overstatement of about four times — and has around a six percent chance of reporting the wrong sign. Adequate power for that effect requires roughly 11,700 per arm.
-
Publication-bias correction can eliminate a headline finding entirely rather than merely shrinking it. Mertens and colleagues (2022, PNAS) reported a pooled nudging effect of about d = 0.43; Maier and colleagues (2022, PNAS) reanalyzed the same data with Robust Bayesian Meta-Analysis and found the evidence for an overall effect essentially vanished, with the Bayes factor supporting the absence of an effect. Szaszi and colleagues argued in the same exchange that a single pooled nudge effect is not a meaningful quantity in the first place. Megastudies (Milkman and colleagues, 2021 and 2022) are the design that fixes both problems, by testing many interventions in one population under one preregistered protocol and reporting every arm.
-
The operating conclusion is to build more measurement capacity, not fewer interventions. Budget one to two percentage points for messaging-class nudges and reserve larger priors for structural changes such as defaults and mandated choice, which alter the outcome of taking no action rather than the framing of information and hold up considerably better — though registration gains from defaults should not be assumed to convert into downstream outcome gains, as the organ donation literature demonstrates. Compute the minimum detectable effect before building, run fewer and larger tests, pool arms against a shared control, and treat any intervention whose realistic effect sits below the programme detection threshold as a judgement call rather than a validated win.
Concepts defined
Read Next
- Behavioral Economics
Decision Fatigue Did Not Replicate: What Survives of Ego Depletion, and What CRO Should Build Instead
Ego depletion did not survive preregistered replication: 23 labs, 2,141 participants, d = 0.04. Many interventions it justified still work, for other reasons. The mechanism determines what a team builds next.
- Behavioral Economics
Peak, End, and Exit: Why Remembered Product Quality Diverges From Experienced Quality
What a customer remembers about a product is not the average of what they experienced and barely reflects how long it lasted. Retention, renewal, and survey scores all run on the remembered version, not the lived one.
- Behavioral Economics
The Goal Gradient Effect: Progress Mechanics, Endowed Advancement, and the Post-Reward Reset
Effort accelerates as perceived distance to a goal shrinks, and the reference point is a product decision. The field evidence is unusually clean. So is the failure mode most teams never measure: the post-reward reset.
The Conversation
Be the first to weigh in
Join the conversation
Disagree, share a counter-example from your own work, or point at research that changes the picture. Comments are moderated, no account required.