Booking.com has stated repeatedly that it runs well over a thousand concurrent A/B experiments on an in-house platform. An experiment ships, every engineering dashboard stays green — latency flat, error rate flat, saturation fine — and bookings for one market fall 4% for eleven days before anyone connects the two. What was missing, and how would you design the metric layer so this is caught in hours?
Show the full answer Hide the answer
The situation they were in
Running experiments in the thousands is not a scaling problem in the obvious sense. The hard part is that at that concurrency, the question "did this change hurt the business?" cannot be answered by anyone looking at a dashboard, because no human is looking at a thousand dashboards and the interactions between experiments are not visible on any one of them.
The green dashboards in the stem are not a failure of monitoring. They are monitoring working correctly on the wrong question. Latency, error rate and saturation describe whether the system is serving requests. They say nothing about whether the requests it serves are producing the outcome the business exists for. A change that renders a disabled button, mis-sorts a results list, or quietly excludes one currency from a filter produces a perfectly healthy service and no revenue.
What the design has to be
Three layers, and the middle one is the one that is missing here.
- System metrics — the four golden signals. Fast, cheap, and they detect breakage.
- Guardrail metrics — a small, fixed set of business outcomes that every experiment is evaluated against, whether or not the experiment's owner thought they were relevant. Conversion rate, revenue per session, bookings per visitor, cancellation rate, support contacts per thousand sessions. The owner of a UI-copy experiment does not pick these; the platform applies them to every experiment automatically. That is the whole mechanism: the person shipping the change is the person least likely to anticipate which business metric it will move.
- Business KPIs — the slow, finance-grade numbers, reconciled daily or weekly. Correct and far too slow to be a detector.
The eleven-day gap is the gap between layer 1 and layer 3, and it exists in every organisation that has not deliberately built layer 2.
Why it fits their constraints
Guardrails work at a thousand concurrent experiments for a reason that would not apply to a dashboard: they are evaluated per-experiment against a control group, not against a time series. A 4% drop in one market is invisible on a global time series — it is inside the seasonal noise — but against the control arm of the experiment that caused it, in that market segment, it is a clean signal with a p-value. The comparison is what makes a small effect detectable, and the control group is what makes it attributable.
The practical consequences to state:
- Automatic stop rules. If a guardrail crosses a pre-set threshold with sufficient power, the platform ramps the experiment down without asking. A human in the loop at a thousand experiments is a human who approves everything.
- Segment the guardrails. A global-only guardrail would have missed this: one market out of dozens, 4%, averages away. Evaluate per market, per device class, per logged-in state. The cost is multiple comparisons, which is handled by requiring a larger effect or correcting the threshold, not by abandoning segmentation.
- Define the minimum detectable effect in advance. "We would notice a 4% drop in a market this size within N days" is an answerable question before the experiment runs, and if the answer is "we would not", the experiment needs more traffic or a longer run, not a better dashboard.
What it cost them
Guardrails are not free, and the costs are the reason teams skip them:
- Instrumentation discipline. Business events must be emitted as reliably as HTTP metrics, with the same alerting on their own pipeline. A guardrail that silently stops receiving data reads as "no change", which is the most dangerous possible failure mode — the detector fails safe-looking.
- False positives. Evaluate a dozen segmented guardrails on a thousand experiments and the platform generates alarms daily by chance alone. This is a statistics problem with known answers, and it is real work, and the alternative is alert fatigue that makes the guardrails ignored.
- Organisational friction. The first time a guardrail halts a senior person's favourite feature, the mechanism is contested. It survives only if the threshold was agreed before the result was known.
When copying this would be the wrong answer
At low traffic, this entire apparatus is worse than useless. A product with 500 sessions a day cannot detect a 4% change in anything within a human timescale; the experiment would need months, during which the product has changed underneath it. The honest version at that scale is to ship, watch the raw numbers, and keep changes small enough to attribute by eye — plus a single end-to-end business check ("did any orders complete in the last hour?") which catches the catastrophic case and is the only thing the volume supports.
The decision rule: guardrail automation is justified when the rate of changes exceeds the rate at which humans can attribute outcomes to them. For most organisations that is dozens of experiments, not thousands, and it arrives earlier than expected.