Behavioral Economics

Peak, End, and Exit: Why Remembered Product Quality Diverges From Experienced Quality

What a customer remembers about a product is not the average of what they experienced and barely reflects how long it lasted. Retention, renewal, and survey scores all run on the remembered version, not the lived one.

Share

TL;DR: What a customer remembers about a product experience is not the average of that experience, and is close to independent of how long it lasted. Retrospective evaluations are dominated by the most intense moment and the final moment, a pattern established in cold-pressor trials (Kahneman, Fredrickson, Schreiber and Redelmeier, 1993), in real-time pain reports from colonoscopy and lithotripsy patients (Redelmeier and Kahneman, 1996), and causally in a 682-patient randomized trial where adding about three minutes of mild extra discomfort at the end made patients rate the procedure as less unpleasant and made them more likely to return (Redelmeier, Katz and Kahneman, 2003). Retention, renewal, referral and survey scores all run on the remembered version. The design implication is that the distribution of moments matters more than the mean, and the most neglected moment in software is the cancellation flow.


The Account That Churned Against Its Own Logs

The situation is familiar enough that most B2B software teams have a version of it in the last four quarters. An account has been on the platform for fourteen months. Weekly active seats are stable and slightly above the cohort median, feature adoption breadth is in the top third, support ticket volume is low. There was one bad week in March when a data import failed, took eleven days to resolve, and required two calls with a solutions engineer. Usage dipped and recovered to trend within a fortnight. Every other week is unremarkable in the way that healthy accounts are unremarkable.

The account does not renew. The exit survey says, in the customer's own words, that the migration was a disaster and they never fully trusted the data afterward.

What happens next inside the vendor is the interesting part. Someone pulls the usage data and observes that March was one week out of roughly sixty, that the incident was resolved, that the customer went on to expand seat count in June, and that the stated reason is therefore not a fair account of what actually happened. The word that gets used is "irrational," or in a more charitable meeting, "recency bias." The conclusion is that the exit survey is unreliable, that this account was an outlier, and that the correct response is to keep doing what the health score says is working.

The customer is remembering correctly

That conclusion is wrong on the most important point. The customer is not misremembering. The customer is remembering correctly, in the sense that their retrospective evaluation is being produced by the process that produces retrospective evaluations in every human being, including the analyst who is calling it irrational. That process does not compute an average. It does not integrate over time. It weights the most intense moment and the final moment heavily and almost everything else lightly, and it is close to blind to duration. Under that process, fourteen months of adequate experience punctuated by one severe negative moment produces exactly the memory the exit survey reported.

This matters commercially because every consequential customer decision runs on the remembered version rather than the experienced one. Renewal is a decision made at a point in time by a person consulting a summary. So is referral. A survey response is, quite literally, a request that the person produce a summary on demand. None of these has access to the moment-by-moment record; each has access only to the reconstruction, which is systematically different from the record in ways that have been characterized for three decades.

The practical consequence is that the instruments most product and customer-success teams rely on are measuring two different objects and treating them as one. Usage telemetry measures the experienced stream. Satisfaction surveys measure the remembered reconstruction. Health scores built primarily on the first are used to predict outcomes driven primarily by the second, which is why they fail in a characteristic direction: green until the renewal conversation, then a surprise. That is not a modeling error better features would fix. It is a category error about what is being predicted.

Two Selves, Two Utilities

The framing that makes this tractable is Kahneman's separation of experienced utility from remembered utility, developed across the body of work discussed below and summarized for a general audience in Thinking, Fast and Slow (2011) as the experiencing self and the remembering self. Experienced utility is the moment-by-moment hedonic state, in principle integrable over the duration of an episode. Remembered utility is the global retrospective evaluation the person produces when asked, after the fact, how it was. These are not two estimates of the same underlying quantity. They are two distinct quantities produced by two distinct processes, and they can be dissociated experimentally.

The reason this distinction has purchase in commercial settings is that the two quantities have different jurisdictions. The experiencing self has the experience, and is therefore what a product designer might naively think they are optimizing. The remembering self decides whether to come back, what to tell colleagues, and what number to put on a survey. Whatever a company believes it is optimizing, the mechanism by which that optimization converts into revenue runs through the remembering self.

Three properties of remembered utility do most of the operational work, and they are worth stating separately because they have different evidentiary status. The peak weighting: the most intense moment of an episode, in whichever direction, carries disproportionate weight. The end weighting: the final moment carries disproportionate weight. Duration neglect: the length of the episode carries very little weight, sometimes indistinguishable from none. The first two are usually bundled as the peak-end rule; the third is a separate finding with a separate literature, and it has the more surprising operational consequences in software.

Two paths from moments to outcomes, and the one most health scores instrument
Loading diagram...

The diagram makes the structural problem visible. Engagement telemetry proxies the integrated experience; the integrated experience has a weak edge into remembered utility; remembered utility drives the decisions that show up in revenue. Health scores are therefore two weak links from the outcome they are built to predict, while the two heavy edges into remembered utility are typically not instrumented at all.

The Cold-Pressor Result

The cleanest demonstration that remembered utility can be dissociated from experienced utility is Kahneman, Fredrickson, Schreiber and Redelmeier (1993), "When More Pain Is Preferred to Less: Adding a Better End," published in Psychological Science 4(6), 401-405. The design is simple enough to state in a paragraph, which is part of why it has survived.

Participants underwent two trials of the cold-pressor task. In the short trial, they held one hand in painfully cold water for sixty seconds. In the long trial, they held a hand in the same painfully cold water for sixty seconds, and then, without being told the protocol had changed, kept the hand submerged for a further thirty seconds during which the water was warmed slightly. The additional thirty seconds were still unpleasant. They were less unpleasant than what preceded them, but they were not comfortable, and they were not neutral.

The arithmetic is not ambiguous. The long trial contains the short trial in its entirety and then adds thirty seconds of genuine discomfort on top. Under any additive account of disutility, the long trial is strictly worse. There is no assumption about the shape of the pain function under which choosing to repeat the long trial makes sense as a minimization of total suffering.

Participants were then asked which of the two trials they would prefer to repeat. A clear majority chose the long one.

Schematic of the Two Cold-Pressor Trials Described in Kahneman et al. 1993 (Illustrative, Not Published Data)

The values plotted above are a schematic reconstruction of the protocol as described in the paper, not the published pain traces, and no quantitative claim should be read off them. What they show is the only thing that matters here: the long trial is the short trial plus an appended segment that is less bad than the segment before it, and the area under the long curve is unambiguously larger.

Two mechanisms are visible at once. The peak-end weighting does the work on the evaluation side, because the long trial ends at a lower intensity than the short trial and its peak is the same. Duration neglect does the work on the accounting side, because the extra thirty seconds of real discomfort did not register as "more." Remove either and the preference reverses.

From the Lab to the Clinic: The 1996 Correlational Study

A laboratory result about ninety seconds of cold water does not automatically apply to anything a company does. The first bridge to a consequential real-world setting is Redelmeier and Kahneman (1996), "Patients' memories of painful medical treatments: real-time and retrospective evaluations of two minimally invasive procedures," in Pain 66(1), 3-8.

Patients undergoing colonoscopy and lithotripsy reported their pain in real time at short, regular intervals throughout the procedure. This produces something almost nobody has for a commercial experience: a moment-by-moment record of experienced utility, collected during the episode rather than reconstructed after it. Afterward, patients gave retrospective global evaluations of how bad the procedure had been.

The retrospective evaluations correlated strongly with the average of the peak pain intensity and the final pain intensity. They were largely uncorrelated with the total duration of the procedure, and duration was not a constant across the sample. Procedure lengths differed by a large multiple between patients. A patient whose procedure lasted several times as long as another patient's, accumulating several times the total reported discomfort, did not on that account remember it as worse.

This is the study that makes the phenomenon feel real to operators, because a colonoscopy is not a laboratory task. It is unpleasant, consequential, and long enough that a duration effect ought to be detectable if duration mattered. It was not detectable.

The Randomized Trial

The study that resolves the ambiguity is Redelmeier, Katz and Kahneman (2003), "Memories of colonoscopy: a randomized trial," published in Pain 104(1-2), 187-194, doi:10.1016/S0304-3959(03)00003-4. Six hundred and eighty-two patients were randomized. In the treatment arm, after the diagnostic portion of the colonoscopy was complete, the colonoscope was left in place, stationary, for about three additional minutes before removal. Leaving a colonoscope in place is not comfortable. It is, however, considerably less uncomfortable than the active portion of the procedure that precedes it.

The treatment arm therefore received strictly more total discomfort than the control arm, structured so that the final segment of the episode was milder than the peak. This is the cold-pressor manipulation, transplanted into a clinical procedure, with random assignment.

Patients in the extended arm rated the overall experience as less unpleasant than patients in the standard arm. They were also subsequently more likely to return for a follow-up procedure.

The second half of that result is what should hold an operator's attention. A change in a self-reported rating is interesting but soft; ratings are cheap and respondents are accommodating. A change in whether a person actually returns for a second unpleasant medical procedure is a behavior with a cost attached. The manipulation moved both, in the same direction, by making the objective experience worse.

Being precise about the epistemic ladder matters, because the three studies are routinely cited interchangeably and do not carry the same weight. The 1993 study demonstrates the phenomenon in a controlled but artificial setting with a within-subject preference as the outcome. The 1996 study shows the same statistical signature in a real clinical setting, but correlationally. The 2003 study licenses a causal claim, in a real setting, with a behavioral endpoint. If you are building an operating recommendation on this literature, 2003 is the load-bearing citation, and the honest version of the recommendation is bounded by what that trial actually manipulated.

The Evidentiary Spine, What Each Study Does And Does Not Establish

StudyDesignWhat It EstablishesWhat It Does Not Establish
Kahneman, Fredrickson, Schreiber and Redelmeier (1993), Psychological Science 4(6)Within-subject cold-pressor. A 60-second trial against a 90-second trial identical for the first 60, then slightly warmedA clear majority preferred to repeat the longer trial, which contained strictly more total discomfort. Peak-end weighting and duration neglect operating togetherThat the pattern extends beyond brief, bounded, affectively intense episodes with a clear terminus
Redelmeier and Kahneman (1996), Pain 66(1)Real-time pain reports at short intervals during colonoscopy and lithotripsy, plus retrospective evaluationsRetrospective ratings track the average of peak and final intensity, not total duration, which varied by a large multiple across patientsCausation. The design cannot rule out that procedures ending badly were worse in ways the peak-end average proxies
Redelmeier, Katz and Kahneman (2003), Pain 104(1-2)Randomized trial, 682 patients. Treatment arm had the scope left stationary for about three extra minutesCausally, that adding mild terminal discomfort improved the retrospective evaluation and raised the rate of return for follow-upThat the same manipulation transfers to voluntary, low-intensity, commercially framed episodes the customer can walk out of
Fredrickson and Kahneman (1993), Journal of Personality and Social Psychology 65(1)Continuous and retrospective ratings of film clips of varying durationDuration neglect as a phenomenon distinct from peak-end weightingThat duration is always neglected. Salience and directed attention restore its weight

Duration Neglect and What It Does to Engagement Metrics

Duration neglect deserves separate treatment because it is the half of the finding that product teams have absorbed least and that contradicts their metrics most directly. Fredrickson and Kahneman (1993), "Duration neglect in retrospective evaluations of affective episodes," in the Journal of Personality and Social Psychology 65(1), 45-55, established it as a phenomenon in its own right using affective film clips of varying length, rated continuously during viewing and globally afterward. The retrospective evaluations were well predicted by the moments of peak and terminal affect and poorly predicted by how long the clip ran.

Three widely used software metrics inherit a problem from this.

Time to value. The dominant onboarding metric in B2B software is elapsed time from signup to some defined first-value event. It is a good metric, but the belief that it is a strong driver of remembered onboarding quality is an extrapolation the literature does not support. What the peak-end account predicts is that the memory of onboarding is set by whether there was an unambiguous first success with a legible payoff, and by the state the user was left in when the session closed. An onboarding that reaches value in four minutes and then dumps the user into an empty dashboard with no obvious next action can produce a worse memory than one that took forty minutes and ended on a visible, correct, shareable result. The elapsed minutes are the part the remembering self discards.

Hours of usage and session counts in health scores. Usage volume is a measurement of the experiencing self's activity. It answers whether the product is being used, which is a real and necessary question. It is a considerably weaker proxy for how the relationship will be summarized at renewal, because summarization does not integrate. This is the mechanism behind the specific failure mode described at the top of this essay, and it explains why the failure has a direction. Health scores built on integrated usage will systematically miss accounts whose experienced stream is fine and whose peak and terminal moments are bad, and they will systematically over-flag accounts with a seasonal usage dip and no negative moments at all.

Tenure as a proxy for loyalty. A four-year customer has accumulated more experienced utility than a one-year customer, and teams reason from this to an expectation of greater attachment. Under duration neglect, accumulated tenure is close to inert in the retrospective evaluation; what the four-year customer actually has is more opportunities to have had a severe negative peak. Tenure does predict renewal in most datasets, but the likely mechanism is switching cost, contractual inertia, and organizational embedding, not a warmer memory.

One limit should be stated before the operating recommendations, because it materially constrains them. Duration neglect is not absolute. Duration recovers its weight when it is made salient, when the respondent has a reason to attend to it, or when there is an explicit expectation against which elapsed time is compared. A user told an import will take ten minutes and watching it take four hours has not neglected the duration; the duration has become the event. A procurement team calculating cost per active hour is attending to duration by construction. The finding is about what people spontaneously do when producing a summary, not about a fixed incapacity.

What the Rule Is Not

The peak-end rule is a heuristic description of how retrospective evaluations get produced. It is not a law, it does not have a fixed coefficient, and the confidence with which it is asserted in product-design writing exceeds what the evidence carries. Four qualifications matter enough to state plainly rather than bury.

The evidence base is short, bounded, intense episodes. A ninety-second cold-water trial, a film clip, a colonoscopy. These share properties that a commercial relationship does not have: an unambiguous start, an unambiguous end, a single continuous affective dimension, high intensity relative to the surrounding day, and an experimenter or clinician who controls the terminus. A multi-month SaaS subscription has none of these. It has no natural terminus until it ends, its affective content is low-intensity and multi-dimensional, and it competes for attention with the rest of the customer's working life. Applying the peak-end rule to a two-year enterprise relationship is an extrapolation.

Field replications are more mixed than the popular treatment suggests. The rule reproduces reliably in designs that resemble the original designs, and less reliably in field settings with ambiguous episode boundaries, extended timeframes, or evaluations respondents have reason to construct deliberately rather than intuitively. This is not a scandal about the original work; it is what one should expect from a heuristic operating on a mental representation of an episode. When the episode is well defined the heuristic has a well-defined input. When it is not, the input is whatever the respondent happens to construct.

The segmentation problem is the deep one. Ask a customer how a vendor relationship has been and the answer depends entirely on what they treat as "the relationship." One respondent segments by contract term and evaluates twelve months. Another segments by their own tenure in the role and evaluates three years. Another, prompted by a support email, segments to the last ticket. Each segmentation has a different peak and a different end, and they can produce sharply different ratings from the same person about the same vendor. Operators do not control the segmentation; they can only influence it through what they make salient in the framing of a question, a much weaker lever than the literature is usually read as offering.

The peak-end average is a model fit, not a mechanism. The formula that the retrospective evaluation approximates the mean of peak and end intensity is a compact empirical regularity, and the relative weight of the two terms is not stable across studies or domains. Sometimes the end dominates, sometimes the peak. Treating the fifty-fifty average as a design constant is a misreading of what the underlying result is.

There is no instrument that measures the experience. Every instrument a company owns measures the reconstruction.

Where the Peaks and Ends Actually Are in a Software Product

With the caveats installed, the operational question is the productive one: if remembered quality is set by a small number of moments rather than by the average, which moments are they, and does anyone currently own them.

The candidate list below is not a taxonomy derived from a study. It is the set of moments that recur, in advisory work, as the ones customers name when asked what the relationship was like, which is the right selection criterion given the construct: whatever moment a customer spontaneously names in a retrospective account is, by definition of remembered utility, functioning as their peak or their end. The chart below summarizes how often each shows up as the named moment in exit conversations. The figures are practitioner estimates from anonymized advisory partner data across a small number of B2B software operators, and should be read as orders of magnitude rather than precision figures. They are not benchmarks, the coding of free-text exit responses is judgment-laden, and the sample skews toward mid-market subscription products with annual terms.

Which Moment Departing Customers Name In Exit Conversations (Advisory Partner Estimates, Directional Only)

Two features of that distribution matter before the individual moments. Only five percent of departing customers name no specific moment: churn narratives inside vendors are usually framed as gradual disengagement and slow erosion of value, and departing customers largely do not describe it that way. They describe an event. And the distribution is flat across seven categories, so there is no single moment a company can fix and be done. There is instead a set of candidate moments of which any given account has one or two live, and the task is to know which one is live per account rather than to optimize the average.

Candidate Peak Moments In A Subscription Software Relationship

MomentTypical DirectionWhy It Carries Disproportionate WeightMinimum Instrumentation
First successful action in onboardingPositive peakThe only unambiguously positive high-intensity moment in most first 90 days. Sets the prior for what the product is forTimestamp of first value event plus a single in-context affect question in the same session
First time something goes wrongNegative peakNot the most severe failure but the first one. It is where the customer forms a belief about reliability, and later failures read as confirmationFlag the first error, failed job or support contact per account, and treat it as a distinct class from the Nth
A support interactionEither direction, highest varianceInvolves a human, is unrehearsed, and happens when the customer is already frustrated. The widest spread on the listPer-ticket transactional survey, plus resolution path and handoff count, not just time to close
Billing or invoicing surpriseNegative peakCombines a financial loss with violation of an expectation the customer believed settled. Overages, seat true-ups, auto-renewal at a higher rateAlert on any invoice deviating materially from the prior period before it is sent, not after it is disputed
Failed migration or data importSevere negative peakHigh effort, high stakes, asymmetric. Success is a mild positive; failure is a severe negative that casts doubt on data the customer cannot independently verifyTrack import and migration attempts as first-class objects with success flag and resolution latency

Candidate End Moments, Where The Episode Gets Fixed

MomentTypical DirectionWhy It Carries Disproportionate WeightMinimum Instrumentation
The renewal decisionFraming momentFunctions as both an end and an evaluation checkpoint. Whatever is salient in the preceding weeks gets weighted as if it were the whole termLog every account-facing event in the 60 days before renewal date as a separate cohort variable
The cancellation flowTerminal moment for the whole relationshipFixes the remembered version of everything that preceded it, and therefore win-back probability, referral behavior, and public review behaviorCohort every cancellation by which flow variant it went through, then follow win-back and referral for at least 12 months

Three of the moments in those two tables deserve elaboration, because they are the ones teams most often mis-handle.

The first failure, not the worst failure. Reliability programs are almost universally organized around severity. The peak-end account suggests ordering matters as much as magnitude, and specifically that the first negative event an account experiences does disproportionate work, because it converts an untested assumption about reliability into a belief with an instance attached. Subsequent failures are then interpreted through the belief rather than evaluated fresh. The same incident therefore has a different memory cost depending on where in the account lifecycle it lands, and no standard severity scheme encodes that.

Billing surprises are structurally worse than they look. A billing surprise is not primarily a money event. It is an expectation violation that arrives with a number attached, which makes it both intense and precise, and precision aids retrieval. Overages, seat true-ups, and price increases at auto-renewal are the three that recur, and they are potent because they typically arrive when the customer is not otherwise engaged with the product, so nothing competes with them for the episode. The cheapest available intervention is to move the surprise earlier and make it a notification rather than an invoice line.

The renewal decision is an end, and it is being treated as a start. Most commercial motions treat renewal as the opening of an expansion conversation. Structurally it is the terminus of a bounded episode, the contract term, and the moments immediately preceding it are weighted as if they characterized the whole term. A support escalation three weeks before renewal is not one incident among sixty weeks; it is the end of the story. Teams that stage their most demanding asks in the pre-renewal window because attention is highest there are trading a little expansion pipeline for a material degradation of the terminal moment.

The Cancellation Flow Is the End of the Story

The most neglected moment on the list is the last one, and the neglect is structural rather than accidental.

Cancellation flows are designed against a single objective function almost everywhere: minimize the number of customers who complete a cancellation once they have entered the flow. This produces the familiar apparatus. The metric that justifies all of it is the save rate, and the save rate is genuinely measurable within the quarter.

Under the remembered-utility account this is close to the worst available design, because it converts the terminal moment of the entire relationship into a negative peak. Everything over the preceding term is then summarized by a person whose final, most vivid, most retrievable memory of the vendor is of being obstructed while trying to leave. That summary determines win-back probability if circumstances change, what they write on a review site, what they say when a peer asks for a recommendation, and what they do at their next company when asked to pick a vendor.

The economics of the trade are not hard to state. If a flow saves a fraction of those who enter it and degrades the remembered relationship for the remainder, the comparison is between revenue booked this quarter from the saved cohort and a set of deferred, diffuse losses in the departing cohort. Firms make this trade badly not because they weight the two terms incorrectly, but because only one of them appears anywhere in the reporting stack. Save rate is a number on a dashboard next week. Win-back rate twelve months out, referral behavior from departed users, and the review-site consequence are usually not measured at all, and a term that is not measured is treated as zero.

Why the cancellation trade is made badly, only one branch is measured
Loading diagram...

It should be said clearly that I know of no clean published randomized trial of cancellation-flow design with win-back or referral as the endpoint. The 2003 colonoscopy trial is the closest proof of concept for the general principle, and it is a colonoscopy. What follows is therefore a reasoned extrapolation supported by mechanism and by consistent but non-experimental practitioner observation, not an established result. It should be treated as a hypothesis with a testable endpoint, which is fortunate, because it is one of the more straightforwardly testable things in a subscription business.

The design that follows from the mechanism is unglamorous. Confirm the cancellation immediately and without a second confirmation screen. Export the customer's data without being asked. Prorate honestly and say so. State the reactivation path in one sentence, including whether the data will still be there. Ask one question, not five, and make it optional. Send one closing message that does not attempt a save. The objective is that the last thing the customer experiences from the vendor is competent and unresentful, because that segment will be doing the retrospective work for years.

Survey Instrumentation: What NPS and CSAT Are Actually Sampling

Net Promoter Score and customer satisfaction scores are retrospective global evaluations. That is not an analogy. It is a literal description: a respondent is asked, after the fact, to produce a summary judgment of an episode. This is precisely the object that the peak-end literature is about, which means these instruments are subject to the entire set of distortions described above, by construction and not by accident.

The consequence teams most often miss is that survey timing does not merely add noise. It changes what is being measured, because it changes which episode the respondent constructs.

A transactional CSAT fired immediately after a support ticket resolves is measuring the ticket. The episode in the respondent's head starts when they contacted support and ends when the case closed. Its peak is the worst moment of the interaction and its end is the resolution, which is usually the best moment, because the problem has just been fixed. This instrument therefore runs high, systematically, for reasons that have nothing to do with the product.

A relationship NPS fired at an arbitrary point in the subscription asks the respondent to construct an episode on the spot with no cue about its boundaries. What they construct depends on what is retrievable, and retrievability favors the recent and the extreme. Relationship NPS is therefore peak-end weighted with a recency-determined definition of "end," where the end is often just whatever the customer last dealt with, including things that have nothing to do with the vendor.

A post-renewal survey measures the term as an episode with the renewal as its terminus, which is the closest any common instrument gets to the object a vendor cares about, and it is administered to a population that has just been selected on having renewed.

None of these is the experience, and there is no instrument in common commercial use that is. The closest available approach is experience sampling, which is what the 1996 study did with its real-time reports at short intervals, and almost no software company runs anything like it. Lightweight in-context micro-surveys at defined moments are the practical approximation, and they are rare because they conflict with the instinct to consolidate all customer feedback into one comparable score.

Indexed Score Level By Survey Trigger Point, Same Operator, Same Quarter (Advisory Partner Estimates, Directional Only)

The figures above are indexed to the in-app random-session prompt at 100 and are practitioner estimates from a small number of advisory partner operators that happened to run more than one trigger concurrently. The point is structural rather than numerical. Instruments that fire at a favorable terminus read materially higher than instruments that fire at an arbitrary point, and the gap is large enough to swamp most real changes in product quality.

This produces a specific and common analytical failure. When a lifecycle or growth team changes a survey trigger, usually to improve response rate, the trend line moves. Going from a seven-day-delayed post-ticket survey to a same-day one, or from an in-app prompt to a post-renewal email, produces a level shift that looks exactly like a change in sentiment. Timing changes are rarely annotated in the BI layer, are typically owned by a different team than the one reading the trend, and are almost never on the list of candidate explanations when someone asks why the score moved. In advisory work a meaningful share of unexplained NPS movements turn out to be instrumentation changes rather than sentiment changes; I would not put a number on the share, because the ones that get investigated are a biased sample of the ones that occur.

The Allocation Question

All of this cuts against the standard approach to experience improvement, which is worth stating in its strongest form first.

The transfer to a commercial setting needs to be made carefully, because the differences between a colonoscopy and a checkout flow are exactly the differences that matter. The mechanism that made the extra three minutes work was that the appended segment was genuinely milder than the segment it followed, in an episode the subject could not exit. Adding a step to a checkout flow reproduces none of those conditions, and the most likely result is abandonment, which is an attrition effect the memory literature has nothing to say about.

What transfers is narrower and still useful: end the episode on its least-bad segment rather than its worst, and where the ordering of segments is a free variable, order them so the hardest part is not last. That is a reordering recommendation, not an addition recommendation. Anyone reading the 2003 trial as a license to add friction to a commercial flow has read past the boundary conditions.

The reordering version is cheap to test and is where I would start. Onboarding sequences, configuration wizards, data-import flows, and renewal paperwork all contain segments whose order is a design choice rather than a technical constraint, and in most of them the hardest step is last because that is how they accreted. Moving it to the middle and closing on a confirmation of something that worked costs a sprint and changes the terminus.

A Framework for Finding the Peak and the End Empirically

The failure mode in applying any of this is assuming you know where the peak and the end are. Teams almost always guess a moment they already own and instrument that. The following sequence is the one that has worked as a starting structure in advisory engagements, and its first step is the one most often skipped.

  1. Choose the episode boundary explicitly, and record that you chose it. Peak-end analysis is undefined without an episode. Session, onboarding, ticket, contract term, whole relationship: each is a legitimate choice and each produces a different peak and a different end. Most teams never make the choice, which is why their conclusions do not compose across analyses. Write the boundary down before collecting anything.

  2. Instrument moment-level affect, not just moment-level events. The telemetry records that an import failed. It does not record how bad that was. A single in-context question at a small number of defined moments, sampled sparsely enough not to become a nuisance, gives you a crude experienced-utility trace. This is not another relationship NPS and should not be reported alongside one.

  3. Recover the peak from the retrospective rather than from your process map. Code free-text exit interviews, cancellation reasons, and renewal-call notes for which moment the customer names. Whatever moment they name is functioning as their peak or end, by definition of the construct. This inverts the usual direction of analysis and is the single highest-yield step, because it routinely surfaces moments that are not on anyone's journey map, invoicing being the most common.

  4. Build the moment-to-outcome regression. Per account, code four variables: intensity of the worst moment, intensity of the terminal moment, an integrated experience proxy from usage, and episode duration. Regress renewal, referral, and relationship survey score on all four. The informative outcome is not confirmation of the peak-end weighting but the coefficient on duration. If duration carries weight in your setting, the boundary conditions are biting and the extrapolation from the clinical literature is weaker for you than the mechanism suggests, which is worth knowing before reallocating a budget on the strength of it.

  5. Test reordering before testing reduction. Hold total friction constant, move the worst segment away from the terminus, and measure. This is a cleaner test of the mechanism than a friction-reduction test, because friction reduction confounds the memory effect with an attrition effect.

  6. Instrument the exit as a first-class cohort. Cohort every departing account by the cancellation flow variant it experienced, survey at thirty and ninety days after departure when the respondent has no commercial relationship and little reason to be polite, and track reactivation and referral by cohort for at least twelve months. Until this exists, the cancellation-flow trade is being made with one of its two terms set to zero.

  7. Annotate every survey instrumentation change. Trigger, delay, channel, wording, sampling frame. Treat each as a dated level shift and segment any trend crossing it.

Underneath the framework is a choice about which self a company is optimizing for, and the two answers imply different products. Optimizing for the experiencing self means minimizing integrated friction, which is a defensible ethical position and a poor revenue strategy under this literature. Optimizing for the remembering self means shaping a small number of moments, which is a better revenue strategy and shades quickly into manipulation if the moments are shaped rather than improved. The 2003 result sits uncomfortably on that boundary: the patients genuinely were made to suffer more, and they genuinely did return for a procedure that was in their medical interest. Whether the analogous move in a commercial setting serves the customer depends on whether the thing they are brought back to is good for them, which is a question the measurement cannot answer and the operator has to.

Key Takeaways

  1. Remembered utility and experienced utility are distinct quantities produced by distinct processes. Retrospective evaluations are heavily weighted toward the most intense moment and the final moment of an episode and are close to insensitive to duration, which means an account with a good average experience and one severe negative moment is predicted to churn, not anomalous when it does.

  2. The evidentiary ladder matters and the studies are not interchangeable. Kahneman, Fredrickson, Schreiber and Redelmeier (1993) demonstrated the preference reversal in a controlled cold-pressor design; Redelmeier and Kahneman (1996) found the same signature in real clinical procedures but correlationally; Redelmeier, Katz and Kahneman (2003) is the 682-patient randomized trial that licenses the causal claim and moved a behavioral endpoint. Build recommendations on the 2003 result and its boundary conditions.

  3. The rule is a heuristic description established on short, bounded, affectively intense episodes with clear termini. Extending it to a multi-month subscription is an extrapolation, field replications outside those conditions are more mixed than the popular treatment admits, and duration recovers its weight whenever it is made salient or the respondent has reason to attend to it. Say so when using it.

  4. Health scores built on integrated engagement fail in a predictable direction: over-optimistic on accounts that had a severe discrete negative moment and recovered usage, over-pessimistic on quiet accounts with clean records. NPS and CSAT measure a reconstruction rather than an experience, so survey timing determines what is measured, and an unannotated trigger change produces a level shift easily mistaken for a sentiment change.

  5. The cancellation flow is the terminal moment of the entire relationship and is almost universally optimized against save rate alone, because save rate is the only term in the trade that appears on a dashboard. The mechanism predicts that a clean, fast, unresentful exit improves win-back and referral, but I know of no published randomized test of this, and the honest status of the recommendation is a testable hypothesis with an obvious experimental design rather than an established finding.

The Conversation

Be the first to weigh in

Join the conversation

Disagree, share a counter-example from your own work, or point at research that changes the picture. Comments are moderated, no account required.

Read Next