A travel platform runs hundreds of concurrent A/B experiments as standing practice; Booking.com has built its product development around very large numbers of simultaneous experiments for years. An architecture team tests a new read-through caching layer in front of the availability service by routing 50% of sessions through it. After nine days the treatment arm shows 1.4% higher conversion at p below 0.01. Both arms are served by the same cache cluster and the same application fleet. What actually happens to the measurement and what would have to be true for the result to mean what the team thinks it means?
Show the full answer Hide the answer
What happens, step by step
Minute one: treatment sessions issue reads through the new layer and populate the shared cache. Minute two: control sessions hit entries that treatment paid to warm, so control gets faster too. The measured difference between arms now understates the true effect, because the control arm has partly received the treatment.
Running the other way, treatment's extra fill traffic competes for the same connections and the same memory. Eviction pressure rises, control's hit rate falls, and control gets worse — so the measured difference overstates the effect. Both leaks are present at once, their sizes are unknown, and the data cannot tell you which dominates. The p-value is computed under an assumption of independent units that the design broke on day one, so it measures precision about a quantity that is not the effect.
Where it amplifies
Any shared mutable resource couples the arms: a cache, a connection pool, a thread pool, a message queue, a search index, an autoscaling group that sizes on total load, a model retrained on pooled traffic. Autoscaling is the nastiest, because it makes the arms' performance a function of the sum of their behaviour, so a treatment that reduces load improves control's latency and looks like no effect at all.
What stops it
Randomise at the level at which the resource is shared. For a change behind a cache cluster that means a cluster-randomised or switchback design: whole caches, shards, data centres or markets assigned to arms, or the change switched on and off across the fleet on a schedule so each unit is compared with itself over time. A 100% rollout against a withheld market works too, with the usual caveat that markets differ.
What it costs
Sensitivity. Session randomisation gives millions of independent units, and a well-run platform can resolve conversion effects of roughly 0.1% to 0.3%. With eight cache clusters as units, the effective sample size is eight, and the minimum detectable effect moves to something nearer 1% to 2% — large enough that most genuine infrastructure wins are invisible. Switchback designs recover some power by using many time blocks, at the cost of carry-over between blocks.
For infrastructure work that is usually the right trade anyway, because the honest alternative is to measure the mechanism rather than the outcome: origin requests per session, p99 at the availability call, error rate, cost per thousand searches. Those are measurable per session without interference, and a conversion claim can then be made from a previously measured latency-to-conversion relationship rather than from this experiment.
When not to randomise by cluster
When the change touches only code on that session's path and no shared mutable state — a different ranking formula computed per request, a copy change, a new layout — session randomisation is correct and far more sensitive, so do not pay for cluster assignment. The test is mechanical: ask what state the treatment writes that the control reads. If the answer is anything, the unit of randomisation is wrong.
The failure this prevents is specific and expensive: a caching change that is actually neutral ships on a contaminated 1.4% reading, the cache becomes a dependency of the availability path, and the next origin incident now degrades every arm at once.