concept

Infrastructure Experiment Confound

An A/B result that cannot be believed because treatment and control share the resource the change alters - so each arm changes the other's behaviour and the measured effect has no reliable size or sign.

experimentationinterferencerandomisationmeasurementcaching

An architecture team ships a read-through cache in front of an availability service, routes 50% of sessions through it, and after nine days reports 1.4% higher conversion at p below 0.01. Both arms are served by the same cache cluster and the same fleet. The number is not a measurement of the change; it is a measurement of two arms contaminating each other, and the data cannot tell you which direction dominates.

Treatment warms shared entries that control then reads, so control partly receives the treatment and the gap understates. Treatment's fill traffic also raises eviction pressure and can slow control, so the gap overstates. The statistics assume independent units, and that assumption was broken on day one.

Why it matters

Infrastructure changes are the ones architects are asked to justify, and the standard measurement tool silently stops working exactly where the change is most architectural. A cache, a connection pool, a queue, a search index, a shared model or an autoscaling group all couple the arms. Autoscaling is the worst case: capacity sizes on the sum of both arms, so a treatment that reduces load improves the control arm's latency and can read as no effect at all.

The consequence is a decision made on a number with an unknown sign — a neutral change shipped as a win, or a real win abandoned.

Implementation patterns

  • Randomise at the level at which the resource is shared: whole cache clusters, shards, data centres or markets as units.
  • Switchback designs turn the change on and off across the whole fleet on a schedule, comparing each unit with itself over many time blocks, which recovers statistical power at the cost of carry-over between blocks.
  • A withheld market or cell run at 0% while the rest runs at 100%, accepting that markets differ and the comparison needs a pre-period baseline.
  • Measure the mechanism per session instead of the outcome: origin requests per session, p99 at the dependency, error rate, cost per thousand searches. These are not confounded, and a business claim can be made from a previously measured latency-to-conversion relationship.
  • Write the sharing analysis into the experiment design review: what state does treatment write that control reads.

Industry example

Booking.com has built its product development around running very large numbers of concurrent A/B experiments for years, and the published figures for concurrency come from case-study work rather than from the company's own engineering posts, so the defensible statement is the practice rather than a number. What matters architecturally is what that practice demands: with hundreds of live experiments, the platform has to manage interference between experiments as well as within one, which is why mature platforms hold assignment units, exclusivity groups and guardrail metrics as first-class platform concerns rather than as analyst conventions.

Failure scenarios

  • A cache change shipped on a contaminated reading, making the cache a dependency of the request path so the next origin incident degrades every arm at once.
  • A connection-pool tuning test where the treatment starves the control of the same pool, producing a large apparent win.
  • Cluster randomisation adopted without recomputing power, so a genuine 0.4% effect is reported as "no significant difference" from eight units.
  • An experiment that moved a bottleneck rather than removing it: the measured step gets faster, the next step saturates, and the outcome metric stays flat.

Trade-offs

Correct design costs sensitivity. Session randomisation gives millions of independent units and can resolve conversion effects near 0.1 to 0.3%. Eight cache clusters give eight units and a minimum detectable effect nearer 1 to 2%, which hides most genuine infrastructure wins. Switchbacks need weeks of calendar time, and withheld markets spend revenue.

For most infrastructure work the right answer is to stop trying to measure conversion directly and measure the mechanism, accepting a weaker causal claim in exchange for a believable one.

When not to use it

Session-level randomisation remains correct and far more sensitive when the change touches no shared mutable state: a ranking formula computed per request, a copy change, a layout, a new onboarding step. Do not pay for cluster assignment there. The test is mechanical — ask what state the treatment writes that the control reads, and if the answer is anything at all, the unit of randomisation is wrong.

Interview question

Q: Your team wants to prove that a new caching layer improved conversion. Design the experiment, then tell me what you would report if the business will not fund a design that takes six weeks.

What a strong answer covers: identify the shared resource and the resulting interference; propose cluster randomisation or a switchback with the power cost stated; offer the fallback of per-session mechanism metrics plus a previously estimated latency-to-conversion relationship; and state plainly that a session-randomised conversion number from a shared cache is not evidence, whatever its p-value.

Quick check

Quiz: Why can a session-randomised A/B test of a shared cache not be trusted? Because the arms share the resource being changed, so each alters the other's behaviour and the measured effect has an unknown size and sign.

Flashcard: What is the one-line test for whether your randomisation unit is valid? Ask what state the treatment writes that the control reads — any shared mutable state means the unit must be the shared resource, not the session.