A retailer with 40 million loyalty members and a brand with 12 million exposed device identifiers want to measure in-store purchase lift inside a clean room. Hashed-email match comes back at about 23%. The four-week category purchase rate is around 2%. Roughly what relative lift can this measurement detect, and what does the number rule out?
Show the full answer Hide the answer
The assumptions, stated
- Matched exposed cohort: 12M exposed identifiers at a 23% match gives roughly 2.8 million people the room can see on both sides.
- A comparable matched unexposed control of similar size, which the retailer can supply because its loyalty base is far larger than the matched set.
- Baseline conversion p = 0.02 over four weeks, measured on the retailer's side.
- 80% power, 5% two-sided significance, one pre-registered primary metric.
The arithmetic, shown
Standard error of a difference in proportions with n per arm is sqrt(2p(1-p)/n).
2 × 0.02 × 0.98 = 0.03920.0392 / 2.8e6 ≈ 1.4e-8, so SE ≈ 1.2 × 10⁻⁴- Minimum detectable effect ≈ 2.8 × SE ≈ 3.4 × 10⁻⁴ absolute
- Relative to a 2% base: about 1.7%
So the design detects a relative lift under 2% on the headline number. That is far more sensitive than the effect anyone is looking for — brand campaigns are argued about at 5% to 20% relative.
Now slice it. Forty reported segments drop n to about 69,000 per arm: SE ≈ 7.5 × 10⁻⁴ and the MDE rises to roughly 10% relative. Push to store-by-week-by-segment and you have 1,200 stores × 4 weeks × 40 segments ≈ 192,000 cells averaging 14 matched people each, every one of them below a minimum-cohort threshold of 1,000 and therefore suppressed.
Which assumption dominates the error
Not n. The 23% match rate dominates, and it does so as bias rather than as noise. A hashed-email match succeeds for people who are logged in, loyalty-enrolled and shop often, and those people have a baseline purchase rate well above the 2% category average. The measured lift describes the matched cohort, and the matched cohort is the retailer's best customers.
The second-order version is worse: if the control group is drawn from the matched population by the same join, the matching itself is correlated with the treatment, because exposure on the brand's side also requires an identified user. Confirm that the control is matched by the same mechanism before believing any number the room returns.
What the number rules in or out
- Rules out "the result was null because the sample was too small". At 2.8 million per arm it was not. Expect that explanation anyway and have the MDE ready.
- Rules out publishing a 40-segment store-level breakdown. The cells are either suppressed or statistically meaningless, and running them anyway accumulates overlapping queries that make the room easier to attack by differencing.
- Rules in spending the next quarter on match quality rather than on query volume: moving match from 23% to 45% changes who the answer is about, which is worth more than any extra precision.
- Rules in pre-registering one primary metric and at most three to five secondary cuts, because the power budget is spent on multiplicity long before it is spent on n.
When this is the wrong answer
When a geographic experiment is available, skip the clean room. Randomise 1,200 stores into exposed and holdout markets, measure at store-week level, and you need no identity join, no match rate and no cross-party data movement at all. The precision is coarser and the estimate is unbiased with respect to who is identifiable, which is the error that actually matters here. Prefer the geo design unless the treatment cannot be assigned geographically — a personalised audience, say — or the brand will not accept a holdout large enough for a geo design.
Retail-media clean rooms have been in production since roughly 2021, and match rates on hashed contact fields in this class of arrangement commonly sit well below half. This failure mode is never flagged by the room: suppression rules, query review and aggregation thresholds all pass, so the only control is the discipline of quoting the match rate beside the result.