intermediate 3 min answer

Review this. Forty live experiments. Each feature calls a central assignment service synchronously when it needs a variant - up to 7 calls on one page render at roughly 9 ms each - and each feature emits its own exposure metric with its own field names. What would you change, what would you leave alone, and what breaks when that service is slow?

booking.comexperimentationcross-cutting concernsdeterminismmetrics
Show the full answer Hide the answer

What is actually required

Three properties, and only three. Assignment must be deterministic and stable for a unit over the life of an experiment. Exposure must be recorded at the moment the variant could have affected the user, not when it was computed. Metrics must be comparable across experiments, which means one schema, not forty.

Everything else in the current design is a choice, and two of the choices are wrong.

What I would change

Compute the whole assignment map once per request, locally, from a deterministic hash of the unit identifier and the experiment salt. Assignment is a pure function of the unit identifier and the experiment configuration, so the network call buys nothing except configuration freshness, and a 30-second local cache of the configuration buys that instead. Removing 7 hops at 9 ms takes about 63 ms off the render and removes a hard runtime dependency from every page.

The subtler gain is correctness. Two calls inside one request can straddle a configuration change and return different variants for the same user, which contaminates the experiment with no error and no log line; the only symptom is a result that will not replicate. Computing once per request makes that impossible by construction.

Second change: one exposure event schema, owned centrally, with the same unit identifier, experiment identifier and variant field names for all forty experiments. Forty naming conventions is forty analyses that cannot be compared or audited.

What I would leave alone

The decentralised ownership of experiment definitions, and emission at the point of use. Only the feature knows the moment the variant mattered, so emission belongs there. Booking.com's 2017 paper Democratizing online controlled experiments at Booking.com describes members of its departments running and analysing more than a thousand concurrent experiments, which is the evidence that decentralised ownership scales. What does not scale is a decentralised schema.

What happens when the assignment service is slow

The usual default is fail-open to control. Nothing errors, latency recovers, and every user served during the timeout window is silently counted as control, biasing results toward no effect for as long as the degradation lasts. Nobody sees it, because the dashboard that would show it is the experiment itself.

With local assignment the failure mode changes shape: a stale configuration means a newly started experiment begins a minute late on some hosts. That is visible, bounded, and harmless.

How I would argue this in the review

With two numbers and one story: 63 ms of avoidable render latency on the critical path, and one named experiment whose result nobody could reproduce. Avoid arguing architecture in the abstract. The team that owns a contaminated result will agree faster than the team that owns the latency.

When this is the wrong change

If a kill switch has to take effect inside the current page view, a 30-second configuration cache is wrong and the synchronous check is the price of that guarantee. Scope it to the one or two experiments that need it rather than all forty. The same applies when a variant depends on a server-side segment the caller does not hold: the call is then real work rather than a lookup, and the fix is to make it once at the edge instead of seven times per render.