A board sets an appetite of at most two customer-visible material incidents a year. Engineering ships about 3000 production changes a year with a change failure rate near 15%. Roughly how many material incidents does that imply, and which lever actually closes the gap?
Show the full answer Hide the answer
The assumptions, stated
An appetite over a count is unusable until the count has a definition, so fix one first: material means more than 5 minutes of customer-visible failure above a 1% error rate on a customer-facing path. Write that into the appetite or every number below is arguable.
Then: 3000 changes a year, change failure rate 15%, and an escape rate — the share of failed changes that reach enough customers for long enough to be material. Most organisations have never measured this. Assume 5% to start.
The arithmetic
- 3000 changes × 15% = 450 failed changes a year, roughly 9 a week.
- 450 × 5% = about 22 material incidents a year, against an appetite of 2.
- The gap is 11x, not 10%. This is not a tuning problem.
Now test the levers one at a time, holding the others fixed.
- Cut change failure rate. To reach 2 incidents at a 5% escape rate you need a failure rate near 0.4%. Published delivery research has consistently placed even strong performers in the tens of percent rather than under one percent, so this is not an available number. Rejected.
- Ship less. 400 changes a year at 15% and 5% gives 3 incidents, and it is the wrong answer anyway: fewer, larger changes raise the failure rate per change and lengthen diagnosis, so the product of the three terms usually goes up. Rejected, and worth saying out loud because it is the move boards ask for.
- Cut the escape rate. 450 failures at 0.44% gives 2. That means 24 of every 25 failed changes must be caught or reverted before they are material, which is an engineering programme with known parts: canary with automatic rollback on error-rate delta, cell or tenant partitioning so a failure reaches a fraction of users, feature flags with a kill switch independent of the deploy path, and above all a revert measured in minutes. If the definition of material is 5 minutes, a 3-minute automated rollback converts most escapes into non-events by definition.
Which assumption dominates the error
The 5% escape rate, by a wide margin. It is unmeasured, it could plausibly be 1% or 20%, and it moves the answer by the same factor. So the first action is not a programme, it is instrumentation: classify every failed change in production by whether a customer saw it, for how long, and at what error rate. Three months of that data replaces the whole estimate.
What the number rules in and out
It rules out the conversation the board expects, which is about approval. It rules in blast radius and revert time as the architecture of risk appetite: the appetite is a statement about impact per failure, and impact per failure is a design property, not a discipline property.
When this is the wrong answer
When the appetite is really about a single catastrophic event rather than a count. "No incident may cost more than 2% of annual revenue" cannot be managed by multiplying rates, because the distribution's tail, not its mean, is the whole concern. There you want a threshold rule per decision and a hard stop, and counting incidents will quietly reassure you right up to the event that ends the argument.