Novelty, Primacy and Long-Term Effects
Treatment effects that grow or decay as users learn, how to detect and measure that learning, and the two ways to estimate a long-term effect without waiting years, long-running holdouts and surrogate indices.
Microsoft News replaced the Outlook.com button on its homepage with a Mail app button. Clicks on the button rose 28%, clicks on the adjacent button rose 27%, and overall page clicks rose 4.7%. Day by day, the gap shrank. Users who expected the old destination kept clicking, landed somewhere unexpected, tried the neighbouring button to see what else had changed, and eventually learned. The team stopped the experiment early (Sadeghi et al., 2022, Novelty and Primacy: A Long-Term Estimator for Online Experiments, Technometrics 64(4)). A two-week readout would have shipped confusion as engagement.
Two directions of user learning
A standard experiment assumes a stationary treatment effect. User learning violates that in two directions. Novelty: a change attracts exploration that fades, so the early effect overstates the long-run one. Primacy: users resist or are slow to discover a change, so the early effect understates it. A new ranking model can show primacy as users learn to trust better results; a redesigned layout can show negative primacy that turns positive once habits re-form.
These are distinct from other reasons an effect changes over time: seasonality interacting with treatment, a population of active users that shifts during the test, and interference that builds as more of a network is treated. Diagnosing learning means ruling those out.
Measuring learning: cohorts and post-periods
Hohnhold, O'Brien and Tang built the reference methodology at Google while studying ads blindness, users' learned tendency to ignore ads (Hohnhold, O'Brien and Tang, 2015, Focusing on the Long-term: It's Good for Users and Business, KDD). Two designs isolate learning from the direct effect of the change.
The post-period design exposes a cohort of users to the treatment for months, then switches it back to control. Any remaining difference from a control cohort, measured on identical serving, is pure learning.
The cookie-cookie-day design re-randomises a fresh set of users into treatment each day, so each day's treatment group has no accumulated exposure. Comparing it with a cohort exposed continuously isolates the learned component while treatment is still running.
They modelled learning as exponential, \(U(t) = \alpha\,(1 - e^{-\lambda t})\), and estimated a half-life of roughly 60 days, \(\lambda \approx 0.012\) per day. That rate sets the price of measuring long-term effects. A 90-day study captures \(1 - e^{-0.012 \times 90}\), about two thirds of the eventual learning effect (the paper rounds to 65%), and measuring over the first 14 days of a post-period loses a further 8% to unlearning, so a standard study sees roughly 60% of the true effect. In an experiment that increased mobile ad load, the short-term revenue gain was significant while the long-term estimate, once learning was accounted for, was essentially zero. Findings like this led to launches that cut the ad load on Google's mobile search by 50%.
Sadeghi and colleagues offer a lighter alternative for detection at scale: treat the time pattern of the treatment effect within an ordinary experiment as a difference-in-differences problem, with no special cohort design. It is cheaper and more powerful for testing whether learning exists, and more exposed to other forms of treatment-by-time interaction such as seasonality.
Estimating long-term effects without waiting
Holdouts answer the question directly: keep a small share of users on the old experience for months after launch. They are expensive in exposure, degrade as users churn or switch devices, and measure the launched bundle rather than any single change.
The surrogate index replaces waiting with prediction (Athey, Chetty, Imbens and Kang, 2025, The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely, Review of Economic Studies 93(4)). Using a separate observational dataset where both short-term outcomes \(S\) and the long-term outcome \(Y\) are seen, fit \(h(s, x) = \mathbb{E}[Y \mid S = s, X = x]\). The long-term treatment effect is then the experiment's effect on \(h(S, X)\):
This is valid under the surrogacy condition \(Y \perp T \mid S, X\), meaning the treatment affects the long-term outcome only through the measured short-term outcomes, plus comparability between the observational and experimental samples. In a California job-training experiment, employment rates from the first six quarters recovered the nine-year employment effect, with standard errors 35% smaller than waiting for the long-term outcome itself.
When it breaks
Surrogacy fails exactly where it is needed. A treatment that improves every short-term proxy can still hurt the long-term outcome through a path no proxy measures. Chen, Geng and Jia showed this surrogate paradox can occur even when the treatment raises the surrogate and the surrogate raises the outcome (Chen, Geng and Jia, 2007, Criteria for Surrogate End Points, JRSS-B 69(5)). Ads blindness is a product example: clicks today reduce attention tomorrow.
Late-period data are a selected sample. Users still active in week eight survived seven weeks of treatment. Comparing late-period treated and control actives conditions on a post-treatment outcome, so an apparent fading effect may be differential attrition rather than learning.
Learning curves extrapolate poorly. An exponential fitted to two months of data implies a long-run asymptote the data never reached. Treat extrapolated long-term effects as model-dependent estimates with wide uncertainty, not measurements.
Running longer is not free. Every extra week delays learning from the result and keeps a share of users on a likely-worse variant. The practical compromise, used at both Google and Microsoft, is to reserve long or cohort-based designs for changes where learning is plausible, and to test for learning routinely everywhere else.
7 flashcards for this concept
Click a card to reveal the answer.