At 09:10 an executive dashboard shows yesterday's conversion down 38%. By 09:40 every pipeline is confirmed healthy, row counts are normal and no schema changed. At 10:50 someone finds a pull request merged at 23:38 that changed the semantic layer's definition of a converted session. Fourteen experiments were mid-flight. Which design change most directly prevents a repeat?
Show the full answer Hide the answer
The trigger
A metric definition is code that rewrites history. Changing what counts as a converted session does not alter yesterday's data; it alters the answer to every question ever asked of it. The dashboard recomputed the full series under the new rule, the whole curve stepped down, and the step landed on the most recent point because that is where people look.
Why it propagated
Experiments are the expensive casualty, not the dashboard. A controlled experiment compares two arms through one measuring instrument, and the comparison is only valid while the instrument is fixed. Change the definition mid-flight and both arms move together, so the difference may still look reasonable while the pre-period and post-period are no longer comparable. Nothing errors and no number looks implausible — the experiment simply measures something that was not asked.
Scale decides how often this bites. Booking.com's 2017 CODE@MIT paper on democratising experimentation (Kaufman, Pitchforth and Vermeer) describes more than a thousand concurrent experiments at that company. At that concurrency, a single definition change lands on dozens of running tests, and with even one definition change a week across a few hundred governed metrics the collision is a weekly event rather than a freak one.
Why detection lagged
Every control in the path was satisfied. The change was reviewed, tested, deployed through the pipeline, and correct — the new definition was the better one. There was no failed job, no schema break, no freshness miss and no quality test to violate. The only signal was a plausible-looking 38% move, and the first 30 minutes went on the pipelines, because that is where teams look when a number changes.
Why the other options fail
- Two reviewers. Review catches wrong changes. This change was right. Adding a reviewer slows the semantic layer and leaves the failure mode untouched.
- Dashboard annotation. Genuinely valuable and would have cut the 100-minute investigation to about five. It is a detection improvement, not a prevention: the experiments were already invalid by the time anyone read the annotation.
- Business-hours freeze. This change was merged at 23:38 and a weekly window would have made the step bigger and later, not absent. Freezes move incidents rather than removing them, and they slow every safe change to constrain a rare unsafe one.
- A 20% day-over-day alert. It would have fired at 09:00, roughly when the human noticed, and it fires on real business events too — a holiday, a campaign, an outage upstream. It buys minutes of detection and no prevention.
The structural fix versus the tempting local fix
Treat metric definitions as versioned artefacts with immutable published versions. An experiment records the version at start. The analysis job resolves that version, and if the definition has moved, it either recomputes both arms under the pinned version or refuses to produce a result. The tempting local fix is process — review, freeze, approval — which constrains the change without protecting the consumer that depends on its stability.
The general lesson: anything that compares across time needs the definition pinned, not just documented. Experiments, cohort retention, year-on-year reporting and regulatory submissions all share this property, and all of them fail silently when it is missing.
When this is the wrong answer
At three experiments a quarter, pinning is ceremony. Version pinning needs a metric store that can serve historical versions, an experiment registry that records them, and an analysis job that fails closed — weeks of platform work and a standing operational burden. Below roughly a dozen concurrent experiments, the dashboard annotation plus a rule that definition changes are announced to the experiment owners costs a day and covers the realistic cases. The pinning machinery becomes necessary when nobody can name the experiments that are currently running.