advanced 3 min answer

A payments and ride-hailing platform of Grab's shape replaced its weekly change advisory board with automated gates fourteen months ago. At 23:40 a configuration change reaches production through the emergency path and skips the canary stage. By 23:46 card authorisations are failing in two markets; recovery takes 51 minutes. The review finds that 38% of production changes in the previous quarter used the emergency path, median approval 90 seconds by whoever was on call. What failed?

change-controlbreak-glassdeployment-gatescoverageincentives
Show the full answer Hide the answer

The trigger

A configuration change with no canary. That is the proximate cause and it is the least interesting part: the gate that would have caught it existed and worked.

Why it propagated

A control's strength is its coverage, not its strictness. Gates that apply to 62% of changes provide 62% of the protection, and the missing 38% is not a random sample — it is enriched for exactly the changes made under pressure, at night, by one person.

The mechanism is an incentive gradient, not a discipline problem. If the normal pipeline takes 35 minutes and break-glass takes 90 seconds, every minute of pipeline latency is a subsidy paid to the bypass. Engineers under a deadline take the faster path and tell themselves the change is small. The 38% figure is the pipeline's latency and flakiness measured in human behaviour. Retry-prone tests and slow integration stages show up here long before anyone opens a dashboard about them.

The second-order effect is that the risk function was told change control had been automated, and it had been, for the changes that did not need it most.

Why detection lagged

Break-glass usage was logged in the incident tool while normal changes were recorded in the pipeline, so no single system held the ratio. Nobody had a denominator. Fourteen months is roughly how long it takes for an escape hatch to become load-bearing without anyone deciding it should, because each individual use is defensible.

The structural fix versus the tempting one

The tempting fix is to make break-glass harder: a second approver, a justification field, a director's sign-off. That raises the cost of honesty and produces better-written justifications for the same changes, and in a real incident it delays recovery.

What works:

  1. One path, two speeds. Emergency changes go through the same pipeline with the slow stages skipped, so the artefact, the record and the rollback are identical. The difference is what is deferred, not what is bypassed.
  2. Make the ratio an SLI. Break-glass changes as a share of all changes, reviewed weekly. Above roughly 2% the finding is about the normal path, not about discipline — go and measure pipeline latency and flake rate.
  3. Automatic consequences. Every break-glass use opens a ticket that expires in 48 hours and requires the deferred checks to be run against production. Unclosed tickets block the next break-glass by the same team.
  4. Attack the latency. Cutting the pipeline from 35 minutes to 8 removes most of the reason to bypass it, and is usually cheaper than the governance response.

The general lesson

Ask of every control: what fraction of the population does it actually operate on, and who chooses? A control whose subjects choose whether it applies to them is a suggestion with logging.

When not to build this machinery

A fifteen-engineer team deploying a handful of times a day does not need break-glass telemetry or expiring tickets. A two-person rule and a deploy log give the same assurance at no cost. Build the ratio metric when the number of changes per week exceeds what one person can read, or when an external assessor is going to ask how you know the gate applies to everything.