TL;DR: Decision fatigue is invoked across CRO and UX as settled science. Its scientific basis, ego depletion, is among the most prominent casualties of the replication crisis: a 23-laboratory preregistered replication (Hagger, Chatzisarantis et al., 2016) recovered an effect of approximately d = 0.04, and a second multisite test designed with input from proponents (Vohs, Schmeichel et al., 2021) found effects far too small to support the standard account. The parole-judges study practitioners cite as proof is, at minimum, seriously contested on case-ordering grounds. The awkward part is that many interventions the theory justified still work. Cognitive load, task-switching cost, per-field friction, and selection effects explain them without a depletable resource, and they make different predictions about what to build next. Keeping a wrong mechanism is not harmless. It generates a wrong roadmap.
The Slide Everybody Has Seen
Somewhere in the deck there is a battery icon. It is full at the top of the funnel and nearly empty at checkout, and the caption underneath says something to the effect that every decision drains the user. The recommendation that follows is almost always one of four: reduce the number of form fields, reduce the number of plan options, reduce the number of steps, or move the hardest decision earlier in the flow while the battery is still charged. The slide usually lands well. The recommendations are usually good ones, and the underlying story is one everybody in the room recognizes from their own experience of a bad Tuesday afternoon.
The story has a name in the academic literature. It is called ego depletion, and for roughly two decades it was one of the most productive research programs in social psychology. It generated hundreds of experiments, a dominant textbook framing of self-control, a bestselling trade book, and an entire secondary literature in marketing, medicine, education, and law. It also gave the applied conversion community something it rarely gets: a mechanism. Not a heuristic, not a rule of thumb, but a causal story about what is happening inside the user that explains why simplification works and predicts where else it should work.
That mechanism did not survive contact with preregistered replication. This is not a fringe methodological quibble, and it is not a case of one contrarian lab failing to find an effect. Two large multisite preregistered studies, one of them designed with input from the theory's own proponents, recovered effects indistinguishable from or barely distinguishable from zero. A separate line of meta-analytic work established that the supportive published record was consistent with substantial publication bias. The physiological sub-theory that gave the concept its most vivid form, the claim that self-control burns measurable blood glucose, was implausible on basic metabolic grounds before any of the replication evidence arrived.
The practical situation this leaves is genuinely strange, and it is the reason this essay is worth writing rather than simply noting the retraction and moving on. Most of the interventions the theory justified still work. Shorter forms do convert better. Fewer plan tiers often do increase selection rates. Reducing steps does reduce abandonment. These are among the most reliably reproduced results in applied commercial experimentation, far more reliable than the laboratory effect that was supposed to explain them. A practitioner could reasonably conclude that the whole affair is an academic matter of no operational consequence, keep the interventions, drop the citation, and carry on.
That conclusion is wrong, and the reason it is wrong is the thesis of this essay. The mechanism is not decoration on top of the intervention. The mechanism is what generates the next intervention. A theory is valuable to an operator precisely to the extent that it predicts things not yet tested. If a team believes shorter forms work because they conserve a depletable daily resource, that team will go on to build a roadmap of things that follow from a depletable daily resource: pacing features, budget-aware sequencing, restoration nudges, time-of-day targeting premised on a stock that refills overnight. If instead the team believes shorter forms work because each additional required field is an independent opportunity to abandon, it will build a completely different roadmap: field-level hazard instrumentation, progressive disclosure, deferred collection, and abandonment recovery keyed to specific fields. Both roadmaps start from the same observation. They diverge immediately afterwards, and only one of them is supported.
Why the Theory Was So Attractive
The origin is a single 1998 paper. Baumeister, Bratslavsky, Muraven and Tice published "Ego depletion: Is the active self a limited resource?" in the Journal of Personality and Social Psychology (74(5), 1252-1265), and the design at the center of it is one of the most vividly memorable in the discipline. Participants were seated in a room containing freshly baked chocolate chip cookies and a bowl of radishes. One group was permitted to eat the cookies. The other was required to resist them and eat the radishes instead. Both groups were then given a geometric puzzle that was, unbeknownst to them, unsolvable. The radish group gave up sooner.
The interpretation was that resisting the cookies had consumed some finite quantity of a general-purpose self-regulatory resource, leaving less of it available for the persistence task that followed. Two features of that interpretation made it unusually powerful. The first is that the two tasks had nothing to do with each other. Resisting food and persisting at a spatial puzzle share no content, no skill, and no obvious cognitive machinery. If the first depletes performance on the second, whatever is being depleted must be domain-general, which immediately makes the theory applicable to essentially any sequence of effortful acts. The second is that the resource had a physical analogy ready to hand. Muraven and Baumeister (2000, Psychological Bulletin) developed it explicitly: self-control resembles a muscle, which fatigues with use in the short run and strengthens with training in the long run.
It is worth being fair about why this program was taken seriously, because the temptation after a replication failure is to treat the original researchers as obviously careless and the field as obviously credulous. Neither is accurate. The muscle metaphor is a real theory with real content. It made specific, falsifiable predictions: that depletion effects should be domain-general, that they should show dose-response with the difficulty of the initial task, that they should be reversible with rest, that they should be trainable over time, and that they should have a physiological substrate. Those predictions were tested, and for a long time the published record appeared to support them. By 2010 there were enough studies for Hagger, Wood, Stiff and Chatzisarantis to publish a meta-analysis in Psychological Bulletin (136(4), 495-525) that put the pooled effect at approximately d = 0.62, which is a medium-to-large effect by conventional standards and comfortably large enough to be practically consequential.
The applied community picked this up quickly and, on the whole, responsibly. If self-control is a finite daily resource, then a checkout flow is a sequence of withdrawals from that resource, and every superfluous decision is a withdrawal that buys nothing. That is a clean, actionable, mechanistically grounded argument for simplification, and it arrived at a moment when the conversion optimization field was hungry for exactly that kind of grounding. The problem was never that practitioners misread the literature. For most of the 2000s and early 2010s, they read it correctly. The literature was wrong.
The Replication Evidence, In Order
The evidence against ego depletion arrived in three distinguishable waves, and the order matters because each wave answers an objection raised against the previous one.
Wave one: the published record was not what it appeared to be
Carter and McCullough (2014, Frontiers in Psychology) examined the 2010 meta-analysis and asked a question that is now routine but was then still novel: is the distribution of published effect sizes consistent with what an unbiased sampling of studies would produce? It was not. The pattern of effects across sample sizes showed the signature of small-study bias, in which studies with small samples report systematically larger effects than studies with large ones. That pattern has a benign explanation and a malign one. Benignly, small studies might be run on populations where the effect really is larger. Malignly, small studies with null results are less likely to be written up, submitted, or accepted, so the small-study end of the published distribution is a filtered sample of the lucky half.
Carter, Kofler, Forster and McCullough (2015, Journal of Experimental Psychology: General, 144(4), 796-815) followed with a more comprehensive set of meta-analytic tests using several bias-correction methods. The conclusion was consistent across methods: once corrected for the apparent publication bias, there was little evidence remaining for a depletion effect of consequential size. This did not prove the effect was zero. Bias-correction techniques are themselves contested, they have known failure modes, and a proponent could reasonably argue that the corrections were too aggressive. What the work did establish was that the published record could no longer be treated as a straightforward tally of independent confirmations. The count of supportive studies had stopped being informative.
Wave two: a preregistered multilab test
Hagger, Chatzisarantis and colleagues published "A Multilab Preregistered Replication of the Ego-Depletion Effect" in Perspectives on Psychological Science (2016, 11(4), 546-573) as a Registered Replication Report. Twenty-three laboratories ran a single protocol, agreed in advance, on a combined sample of 2,141 participants. The meta-analytic effect across the 23 labs was approximately d = 0.04.
The design of a Registered Replication Report is what makes this more informative than any single study on either side of the question, and it is worth understanding the format properly because it is the most useful single instrument a practitioner has for evaluating a contested behavioural claim.
Three properties follow from that structure. First, the file-drawer problem is closed by construction: every lab that starts is in the final tally, so there is no filtered sample. Second, the analysis plan is fixed before the data exist, which removes the researcher degrees of freedom that Simmons, Nelson and Simonsohn (2011, Psychological Science) demonstrated are sufficient to manufacture statistically significant support for essentially any hypothesis. Third, because many labs run the same protocol on different populations, the result contains information about heterogeneity as well as about the average, so a genuine effect that only appears under specific conditions has a chance to show up as variance across sites even if the pooled average is small.
The predictable objection to the 2016 result was that the protocol was a bad one. The RRR used a letter-crossing task as the depletion manipulation rather than one of the paradigms with the strongest published track record, and proponents argued, not unreasonably, that a weak manipulation cannot deplete anything and therefore cannot test the theory. This is a legitimate criticism of any replication, and it is the criticism that wave three was designed to answer.
Wave three: a paradigmatic test built with proponent input
Vohs, Schmeichel and colleagues published "A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect" in Psychological Science (2021, 32(10), 1566-1581). The design point of this study was that the protocol was constructed in consultation with researchers who believed the effect was real, using paradigms they endorsed as adequate tests. Once again the study was preregistered and run across many sites. Once again the effects were very small, and not of a size or pattern that supports the standard resource account.
This is the part of the sequence that closes the argument for practical purposes. The first wave showed the existing literature was unreliable. The second showed that a preregistered test found nothing, and drew the objection that the test was inadequate. The third ran a test the proponents themselves helped specify and still found nothing that would support the theory as it had been applied. At that point the reasonable posture is not "the jury is out." It is that the specific claim of a domain-general depletable self-control resource, in the form that was exported into applied practice, is not supported.
The count of supportive studies stopped being informative in 2014. Everything published before that has to be re-read, not simply re-cited.
It is worth saying explicitly what this does not establish. It does not establish that self-control effort has no downstream consequences at all. It does not rule out much smaller effects, effects specific to particular tasks, or effects that operate through motivation and attention allocation rather than through a fuel-like resource. Inzlicht and Schmeichel (2012, Perspectives on Psychological Science, 7(5), 450-463) proposed exactly such a revision, in which apparent depletion reflects a shift in motivation and attention away from further effortful control rather than the exhaustion of a substance. Job, Dweck and Walton (2010, Psychological Science, 21(11), 1686-1693) had already shown that depletion-like effects depended on participants' implicit beliefs about whether willpower is limited, which is difficult to reconcile with a purely metabolic account and much easier to reconcile with a motivational one. Those accounts remain live. What is not live is the version on the slide with the battery.
The Glucose Sub-Theory as a Cautionary Tale
The most instructive episode in the whole sequence is the one that should have been caught first, because it did not require any replication evidence at all to be doubted.
Gailliot and Baumeister (2007, Personality and Social Psychology Review) proposed that the self-control resource was, quite literally, blood glucose. Acts of self-control were said to consume glucose; depleted performance was said to follow from lowered blood glucose; and consuming a sugary drink was said to restore self-control performance. The appeal is obvious. It converts a metaphor into a measurable physiological quantity, it explains why willpower feels like a bodily rather than a purely mental thing, and it comes with an intervention that anyone can run.
The physiological problem with it is not subtle. The brain is metabolically expensive at rest and does not become dramatically more expensive when a person is asked to cross out letters or resist a cookie. Kurzban (2010, Evolutionary Psychology) laid out the argument directly: the marginal glucose cost of a cognitively demanding task, relative to an undemanding one, is far too small to produce the systemic blood glucose changes the theory requires, and far too small for a sugary drink to plausibly work as a restorative in the way described. Whole-brain metabolic rate is remarkably stable across cognitive states. Regional shifts occur, but the total is close to constant, and the circulating pool is regulated on a scale that dwarfs any plausible task-related draw. A related line of work found that merely rinsing the mouth with a carbohydrate solution, without swallowing it, produced effects similar to ingesting it, which is incompatible with a fuel account and compatible with a signalling one.
The applied residue of the glucose theory is still visible. Recommendations to schedule important decisions before lunch, to keep snacks near the sales floor, to time high-friction onboarding steps around meals: all of these trace back to a mechanism that was never physiologically credible. Some of them may still be sensible for entirely different reasons, and that is exactly the pattern this essay is about. A practice can be defensible while its stated justification is nonsense, and when that happens the practice survives only for as long as nobody extrapolates from the justification.
The Parole Judges Study
No single study is cited more often in practitioner writing about decision fatigue than the Israeli parole rulings paper, and none is more frequently presented as though it were settled. It deserves an extended and fair treatment, because the way its evidentiary status has been flattened in retelling is itself the clearest illustration of the problem.
What the original reported
Danziger, Levav and Avnaim-Pesso published "Extraneous factors in judicial decisions" in PNAS (2011, 108(17), 6889-6892). They observed parole board hearings conducted by Israeli judges, who worked through three decision sessions per day separated by food breaks. The reported pattern was that the proportion of favourable rulings began at roughly 65 percent at the start of each session and declined steadily to near zero by the end of that session, then reset to roughly 65 percent after the break. The paper interpreted this as evidence that repeated decision-making degrades judgement and pushes decision-makers toward the status quo default, with the food break restoring capacity.
As a piece of communication this is close to perfect. It has a real-world stake no laboratory study can match, a dramatic result, an intuitive mechanism, and a shocking implication about the fairness of the justice system. It travelled accordingly, into trade books, journalism, law-school curricula, conference talks, and a very large quantity of conversion optimization content, almost always as a settled demonstration that decision quality degrades with decision count.
The critique
Weinshall-Margel and Shapard published a letter in PNAS (2011, 108(42), E833) making a structural point about the data rather than a statistical one. The order in which cases appear within a session is not random. Israeli parole hearings are scheduled according to institutional conventions, and at least two of those conventions correlate with case outcome. Prisoners represented by an attorney tend to be scheduled earlier within a session. Cases originating from the same prison are heard together as a block. Both facts mean that position within a session carries information about the case itself, independent of anything happening inside the judge.
This is not a minor caveat. If represented prisoners are heard first and representation improves outcomes, then a decline in favourable rulings across a session is exactly what one would observe even if the judges were entirely unaffected by the passage of time. The observed association between position and outcome is then a scheduling artefact, and the causal story about depletion is an interpretation laid over a confound rather than an inference from a controlled comparison.
The reanalysis
Glöckner published "The irrational hungry judge effect revisited" in Judgment and Decision Making (2016, 11(6), 601-610) and made a different, and in some ways more damaging, argument. Rather than pointing at a specific confound, he asked whether the reported effect size is plausible for any psychological process at all. The magnitude implied by a decline from roughly 65 percent to near zero within a session is enormous. Psychological state variables of the kind depletion describes do not usually move real-world professional judgement from a majority-favourable rate to a near-zero rate over the course of a few hours. Using simulation, Glöckner showed that plausible case-ordering mechanisms, including rational decision-makers who take on cases as time allows and defer complex cases they cannot finish before a break, can reproduce patterns like the reported one without invoking any fatigue process whatsoever.
The specific mechanism in that class of explanations is worth stating clearly, because it is elegant and because it has an exact analogue in commercial data. If a judge knows a break is approaching and cases vary in expected duration, and if favourable rulings take longer to process than denials, then the cases that get heard immediately before a break will be disproportionately those the judge expects to be quick, which are disproportionately denials. No fatigue is required. The apparent time trend is produced by a scheduling optimisation operating on heterogeneous cases.
The Parole Rulings Evidence Chain, Original Through Reanalysis
| Contribution | Publication | Core Content | Effect on the Practitioner Claim |
|---|---|---|---|
| Original finding | Danziger, Levav and Avnaim-Pesso (2011), PNAS 108(17), 6889-6892 | Favourable parole rulings reported at roughly 65 percent at session start, declining to near zero by session end, resetting after each food break | Cited as direct real-world proof that decision quality degrades with decision count |
| Structural critique | Weinshall-Margel and Shapard (2011), PNAS 108(42), E833 | Case ordering within a session is not random: represented prisoners are scheduled earlier, and cases from the same prison are heard as a block, both correlated with outcome | Position within session carries case information, so the time trend is confounded with case composition |
| Plausibility reanalysis | Glöckner (2016), Judgment and Decision Making 11(6), 601-610 | Simulation showing the reported effect magnitude is implausibly large for any psychological state process, and that ordering artefacts alone can reproduce the pattern | Removes the basis for treating the result as a clean estimate of a fatigue effect |
| Current status | Contested, not retracted | The original data and the reported association stand; the causal interpretation is disputed on ordering and magnitude grounds and has not been resolved by a design that removes the confound | The study cannot carry the weight practitioner citations place on it |
The honest summary is not that the study was debunked. It was not retracted, the data are what they are, and the association between within-session position and ruling is real. The honest summary is that the causal claim rests on an assumption of random case ordering that does not hold, that a plausible non-psychological process reproduces the pattern, and that the effect size is difficult to reconcile with any known psychological mechanism. That is enough to disqualify it as a load-bearing citation. A practitioner who cites it as demonstrated proof of decision fatigue is presenting a contested interpretation of an observational dataset as an established causal fact, which is a category error regardless of how the underlying question eventually resolves.
What Actually Survives
The constructive half. Several well-supported constructs explain most of what practitioners attribute to decision fatigue, and they are not variations on the same idea. They are distinct mechanisms with distinct evidence bases and, critically, distinct predictions.
Cognitive load
Cognitive load theory, developed by John Sweller beginning with his 1988 paper in Cognitive Science (12(2), 257-285), concerns the limited capacity of working memory during a task. Working memory holds a small number of items at once. Cowan (2001, Behavioral and Brain Sciences, 24(1), 87-114) put the practical capacity at around four chunks, revising the older and more famous estimate downward. When a task requires holding more than that, performance degrades: errors rise, completion time rises, and abandonment rises.
The difference between this and depletion is not a detail. Cognitive load is a within-task constraint on how much must be held in mind at once. It applies while the task is being performed and disappears when the demand is removed, and it is relieved by restructuring the task rather than by rest or recovery. It is not a stock that drains over a day and refills overnight. A user who has just completed a demanding form does not have a diminished working memory capacity for the next form; she has exactly the same capacity she had before. That single distinction invalidates most of the roadmap items a depletion belief generates, and it validates a different set: chunking, progressive disclosure, removing the need to hold information across screens, inline validation so errors do not require re-derivation, and defaults that eliminate the need to construct a value from scratch.
Attention and task-switching costs
Switching between tasks is measurably costly. Rogers and Monsell (1995, Journal of Experimental Psychology: General) established the basic switch-cost paradigm, and Monsell (2003, Trends in Cognitive Sciences, 7(3), 134-140) reviewed the mature literature. Switching incurs a reliable time penalty and an error penalty, and part of the cost persists even when the person knows the switch is coming and has time to prepare. This is one of the better-replicated effects in cognitive psychology, and it has direct application: a flow that repeatedly moves the user between different kinds of decision, or between decision-making and information-gathering, imposes real costs that a flow with grouped, homogeneous steps does not.
Again the prediction structure differs from depletion. Switch costs are a function of the number of switches and their dissimilarity, not the total quantity of decisions. Ten homogeneous decisions in a row cost less than five decisions alternating with five lookups. Depletion predicts the opposite ranking, because it counts acts of control rather than transitions between them.
Friction as a hazard, not a fatigue
The most useful reframe for conversion work is the least psychological one. Every additional required field is an additional opportunity to abandon. This needs no mental-state story at all. Each field carries some probability that the user does not have the information, does not want to disclose it, mistypes it and hits a validation error, gets interrupted, or simply reconsiders. Those probabilities compound multiplicatively across fields. The right formal frame is survival analysis: the flow is a sequence of stages, each stage has a hazard rate, and completion is the product of survival across all stages.
This framing is not merely a different vocabulary for the same claim. It makes concrete, testable predictions that a depletion model does not, because it locates the cause in specific fields rather than in accumulated exertion. Under a hazard model, the abandonment cost of a field is a property of that field, and the same field costs the same whether it appears third or eighth. Under a depletion model, the cost of a field is partly a property of its position, because the resource is more depleted later. These predictions diverge sharply and are cheap to test with instrumentation most teams already have.
The chart above is model output, not data. It plots two curves that both end at the same place: 45 percent of starters reach the final field. One is a constant-hazard process in which every field carries the same independent 7 percent chance of abandonment. The other is a budget process in which early fields are nearly free and abandonment accelerates as a notional capacity is consumed. Aggregate completion is identical. The shapes are entirely different, and a team measuring only aggregate completion cannot tell them apart. A team measuring per-field drop-off can tell them apart on the first day of instrumentation, which is why field-level hazard reporting is worth more than almost any other measurement investment in a checkout flow. It is also worth noting that in practice neither curve is what most flows produce: real forms tend to show spiky, field-specific drop-off concentrated at two or three particular questions, which is a third shape that neither idealised model predicts and which points directly at what to fix.
Time-of-day effects are real and mostly compositional
Operational data routinely shows that conversion rates vary by hour of day, and this is often presented as field confirmation of decision fatigue. It is not, because the comparison is not between the same people at different times. It is between different people. Who is on a site at 23:00 differs systematically from who is on it at 11:00 in device mix, traffic source, intent stage, session context, and demographic composition. Late-evening mobile traffic arriving from social referrers is a different population from mid-morning desktop traffic arriving from branded search, and it would convert differently even if human self-control were perfectly constant across the day.
This is a selection effect, and it has a well-understood remedy: compare within-user across times rather than across users at different times, or condition on the observable composition variables and check whether the time coefficient survives. In advisory work the time coefficient usually shrinks substantially once traffic source, device, and returning-versus-new status are controlled. It does not always vanish, and there are genuine circadian effects in the cognitive literature that could contribute. But the default interpretation of an uncontrolled hour-of-day curve should be composition, not fatigue.
Choice overload, which has its own replication problem
One more construct deserves mention because it usually appears on the same slide. The claim that more options reduce choice, popularised by the jam-tasting study, has been meta-analysed. Scheibehenne, Greifeneder and Todd (2010, Journal of Consumer Research, 37(3), 409-425) pooled the available experiments and found a mean effect close to zero with substantial heterogeneity: sometimes more options help, sometimes they hurt, and the moderators determining which are not well established. The operational reading is not that assortment size never matters. It is that "fewer options is better" is not a general law, and therefore has to be tested per context rather than assumed.
Where the Models Disagree
The payoff of all of this is a table. Depletion and the surviving mechanisms are not interchangeable stories about the same thing; they make different predictions, and most of those predictions are cheap to test.
Divergent Predictions From a Depletion Model Versus the Surviving Mechanisms
| Question | Depletion Model Predicts | Surviving Mechanisms Predict | What the Evidence Shows |
|---|---|---|---|
| Does a demanding task impair an unrelated later task? | Yes, because the resource is domain-general | No general effect; only carryover through attention, mood, or motivation, which is task-specific and small | Preregistered multilab tests find effects near zero (Hagger et al. 2016; Vohs et al. 2021) |
| Does the same form field cost more when placed later in the flow? | Yes, position increases abandonment because capacity is lower | No, hazard is a property of the field, not its position, except through cumulative exposure | Testable directly by randomising field order; house recommendation is to run this before believing either |
| Do users have a fixed daily budget of decisions? | Yes, with overnight replenishment | No; working memory limits are within-task and reset when the task ends | Cognitive load theory is a capacity constraint, not a stock (Sweller 1988; Cowan 2001) |
| Does a sugary drink restore self-control performance? | Yes, glucose is the resource | No; the metabolic cost differential is far too small, and mouth-rinse results contradict a fuel account | Physiologically implausible on independent grounds (Kurzban 2010) |
| Do ten homogeneous decisions cost more than five decisions alternating with five lookups? | Yes, because it counts acts of control | No, the alternating sequence costs more because switch costs dominate | Switch costs are among the better-replicated effects in cognitive psychology (Rogers and Monsell 1995; Monsell 2003) |
| Are late-night conversion rates lower because users are depleted? | Yes, capacity is lowest at the end of the day | Mostly no; the late-night population differs in source, device, intent, and tenure | Within-user estimates typically shrink the time coefficient substantially in partner data |
| Does reducing the option count always increase selection? | Yes, each option is a withdrawal | Sometimes; the effect is heterogeneous and condition-dependent | Meta-analytic mean close to zero with high heterogeneity (Scheibehenne, Greifeneder and Todd 2010) |
| Should the hardest decision be moved earliest? | Yes, spend the budget while it is full | Only if the hard decision is also the one users are most committed to completing; otherwise it raises early hazard and loses users who would have converted | Order effects are real but driven by commitment escalation and information availability, not capacity |
Read down the third and fourth columns and the practical stakes become concrete. Four of these eight questions correspond to roadmap items that teams actually build. Budget-aware pacing, restoration nudges, glucose-adjacent timing, and hard-decision-first sequencing all follow naturally from a depletion belief and all rest on predictions the evidence does not support. The last row is the most commercially expensive of them, because moving the hardest step earliest is a common and confidently made recommendation, and under a hazard model it is exactly backwards: front-loading the highest-hazard step maximises the number of users who abandon before accruing any sunk commitment.
Shortening forms does improve completion, reliably enough that it is among the least controversial results in applied conversion work. But the depletion account of why is not supported, and the account a team holds determines what it builds next. A depletion belief predicts, among other things, a fixed daily budget of decisions that replenishes with rest; order effects that carry across entirely unrelated tasks; and remedies based on physiological restoration. None of the three holds up. The daily-budget prediction has no support in the preregistered record. The cross-task carryover prediction was tested directly by two large multisite studies and did not appear. The restoration prediction depends on a metabolic story that was never physiologically credible.
A friction and hazard model, by contrast, makes predictions that are both different and immediately testable with instrumentation most teams already have. It predicts that abandonment is concentrated in specific fields rather than distributed smoothly across position, so a per-field drop-off chart should be spiky rather than a smooth accelerating curve. It predicts that removing one high-hazard field produces a larger gain than removing two low-hazard ones, even though the depletion account scores the second as a bigger reduction in total decisions. It predicts that a field which cannot be removed can still be de-risked by changing its hazard, through defaults, inline validation, format tolerance, or deferral to after the commitment point, without changing the decision count at all. And it predicts that field order matters much less than field identity, which is directly testable by randomising order and which, in partner data, is the test that most reliably surprises teams who came in believing the depletion story.
Which field should you remove?
A hazard model says survival through a form is the product of surviving each field, so one bad field costs more than several ordinary ones. Set the shape of your own form and compare. The defaults are the twelve-field, seven-percent example used above.
Completion rate, percent of starters
completion now
33.8%
remove the worst field
45%
remove two ordinary fields
39%
The comparison is the point rather than the absolute numbers, which depend entirely on inputs you have to measure rather than assume. Removing the single worst field beats removing two ordinary ones across essentially the whole plausible parameter range, and the gap widens as the worst field gets worse. A depletion account scores those two options the other way round, because it counts decisions rather than hazards: two fields removed is two units of budget saved, one is one. That is the divergence in miniature, and it is settled by per-field instrumentation rather than by argument.
The uncomfortable version of this is that a team can run a successful test programme for years on a wrong theory, because the theory happens to recommend a class of interventions that work for other reasons. The cost is invisible: it is the set of experiments never run because the wrong model did not suggest them, and the set of experiments run and failed because the wrong model predicted they would work.
How to Update on a Famous Finding
The last section is about epistemics, because that is the part that transfers to the next famous finding rather than only to this one.
The asymmetry a practitioner most needs to internalise is between a Registered Replication Report and a bestselling book. These are not two sources of comparable authority that happen to disagree. A trade book is a synthesis written at a point in time by an author with a thesis, reviewing a literature whose publication filter was unknown to them, and once printed it does not update. An RRR is a purpose-built instrument for answering one question with the file drawer closed and the analysis plan fixed in advance. When the two conflict, it is not close.
This is not a criticism of the authors of those books, and it is worth saying so directly because the point is often made uncharitably. Daniel Kahneman, whose Thinking, Fast and Slow (2011) discussed both ego depletion and the parole rulings study, later stated publicly that he had placed too much faith in underpowered studies when writing the book's treatment of priming research. That is a model of how to handle the situation, and it should raise rather than lower a reader's confidence in the author. The lesson is not that the books were dishonest. It is that a synthesis of a biased literature inherits the bias, however careful the synthesist, and that the correction arrives later and travels far less widely than the original claim.
The base rate for a published effect
The broader context is the Open Science Collaboration's 2015 project in Science (349(6251)), which attempted direct replications of 100 psychology studies. Ninety-seven of the originals had reported statistically significant results; roughly 36 percent of the replications did. That number is often quoted as an indictment of the field, and it is certainly a serious finding, but for an applied reader the more useful inference is a prior: a randomly selected published psychology effect, absent further information, has a meaningful probability of not being there. That prior should be updated up or down by specific features of the individual claim, which is what a vetting protocol is for.
A Triage Protocol for Behavioural Claims Before They Enter a Roadmap
| Check | What To Look For | Strong Signal | Weak Signal |
|---|---|---|---|
| Replication status | Search the claim plus the words replication, registered report, and multilab | A preregistered multisite replication exists and supports the effect | Only the original and citing papers appear, with no independent direct replication |
| Publication-bias screen | Whether any meta-analysis has applied bias correction, and what happened to the estimate | Corrected and uncorrected estimates are similar | The estimate collapses under correction, or nobody has checked |
| Mechanism plausibility | Whether the proposed mechanism is consistent with better-established adjacent fields | The mechanism is independently documented outside the originating literature | The mechanism exists only inside the behavioural literature that needs it |
| Effect size sanity | Whether the reported magnitude is plausible for the kind of process being described | Effect is modest and consistent with related well-established effects | Effect is dramatic, memorable, and much larger than neighbouring effects on similar outcomes |
| Identification in field studies | For observational field results, whether assignment to the treatment condition is plausibly as-good-as-random | Randomisation, or a credible natural experiment with a defensible control | Ordering, scheduling, or selection could produce the same pattern |
| Retelling distance | How many hops separate the version in the deck from the primary source | The primary source has been read and its critiques searched | The claim arrived via a talk citing a book citing a summary |
| Prediction divergence | Whether the mechanism predicts anything the team could test cheaply and has not | There is a specific, falsifiable, low-cost test the mechanism implies | The mechanism only ever explains results after they arrive |
The last row is the one that does the most work and is the least often applied. A mechanism that only ever explains outcomes retrospectively is not functioning as a theory; it is functioning as a narrative. The test of whether a team actually believes a mechanism, as opposed to using it as post-hoc commentary, is whether the mechanism has ever caused the team to run an experiment it would not otherwise have run, and whether that experiment could have come out the other way. Ego depletion, as it was used in applied practice, almost never met that bar. It was invoked to explain simplification results that had already been obtained, and the roadmap items it uniquely predicted were rarely isolated and tested as tests of the theory.
A mechanism that only ever explains results after they arrive is not a theory. It is commentary with citations attached.
The posture to hold instead
A final point about posture, which is easy to get wrong in both directions. The overcorrection to a replication failure is blanket cynicism: test everything, believe nothing, hold no model. That posture is worse than the one it replaces, because a pure test-everything programme has no way to generate hypotheses worth testing and no way to generalise a result beyond the exact surface it was measured on. A mechanism is a compression of many possible experiments into a small number of commitments. A wrong mechanism compresses badly and generates a bad search. The remedy is a better mechanism, held with appropriate confidence and revised when the evidence moves, not the abandonment of mechanisms altogether.
Ego depletion was a serious theory that made real predictions and turned out to be wrong, which is a normal and respectable outcome for a scientific hypothesis. What was not respectable, and what practitioners share responsibility for, is the decade in which the claim circulated in applied contexts as settled fact while the primary literature was in open dispute. The corrective is not to stop reading behavioural science. It is to read it with the publication filter in mind, to weight preregistered multisite evidence above everything else, to check mechanisms against adjacent fields, and to notice when a claim is being asked to carry more weight than its design can bear.
Key Takeaways
- Ego depletion, the theory behind practitioner claims about decision fatigue, is not supported by preregistered evidence. The 23-laboratory Registered Replication Report of Hagger, Chatzisarantis et al. (2016), covering 2,141 participants, produced a pooled effect of approximately d = 0.04, and the multisite paradigmatic test of Vohs, Schmeichel et al. (2021), designed with proponent input, again found effects too small to support the standard account.
- The published record that appeared to support the theory was compromised before either replication was run. Carter and McCullough (2014) and Carter, Kofler, Forster and McCullough (2015) showed the evidence base was consistent with substantial publication bias, and that little support survived bias correction. This means the count of supportive studies published before roughly 2014 is not informative on its own.
- The parole-rulings study of Danziger, Levav and Avnaim-Pesso (2011), the most-cited real-world evidence in practitioner writing, is seriously contested. Weinshall-Margel and Shapard (2011) showed case ordering within sessions is non-random in ways that correlate with outcome, and Glöckner (2016) argued via simulation that the reported effect is implausibly large and reproducible by ordering artefacts alone. The study has not been retracted and the association is real, but the causal interpretation cannot bear the weight practitioner citations place on it.
- Most interventions justified by decision fatigue still work, for other reasons. Cognitive load is a within-task working memory constraint rather than a daily budget; task-switching costs depend on the number of transitions rather than the number of decisions; per-field abandonment is better modelled as a hazard process than as fatigue; and time-of-day conversion patterns are largely compositional, reflecting who is present rather than what time it is.
- The mechanism matters because it determines the next experiment, not the last one. Depletion uniquely predicts a replenishing daily budget, carryover across unrelated tasks, glucose-based restoration, and hardest-decision-first sequencing, none of which hold up. A hazard model predicts spiky field-level abandonment, larger gains from removing one high-hazard field than two low-hazard ones, and order effects driven by field identity rather than position. Those predictions are cheap to test, and per-field hazard instrumentation is the single measurement that separates the two models fastest.
Tags
Concepts defined
Read Next
- Behavioral Economics
Peak, End, and Exit: Why Remembered Product Quality Diverges From Experienced Quality
What a customer remembers about a product is not the average of what they experienced and barely reflects how long it lasted. Retention, renewal, and survey scores all run on the remembered version, not the lived one.
- Behavioral Economics
The Goal Gradient Effect: Progress Mechanics, Endowed Advancement, and the Post-Reward Reset
Effort accelerates as perceived distance to a goal shrinks, and the reference point is a product decision. The field evidence is unusually clean. So is the failure mode most teams never measure: the post-reward reset.
- Behavioral Economics
Choice Overload Is a Conditional Effect: Four Moderators That Decide Whether Cutting Options Works
The jam study made too much choice a folk theorem. The meta-analytic mean effect is approximately zero. Four moderators decide whether an assortment cut helps or destroys the tail.
The Conversation
Be the first to weigh in
Join the conversation
Disagree, share a counter-example from your own work, or point at research that changes the picture. Comments are moderated, no account required.