Change Risk Budget
also called Material Incident Budget, Change Risk Allowance
A stated ceiling on material customer-visible incidents per period, converted into the change failure rate and escape rate engineering must design for - a conversion that usually shows blast radius and revert time are the only available levers.
A board sets an appetite of at most two material customer-visible incidents a year. Engineering ships about 3000 production changes a year with a change failure rate near 15%. That is 450 failed changes, and if one in twenty reaches customers materially, about 22 incidents against an appetite of 2. The gap is eleven times, and no approval process closes a gap of that size.
A change risk budget is the appetite written so the arithmetic can be done. Until it is converted the count is a sentiment, and engineering keeps being asked for approval when the problem is impact per failure.
Why it matters
Boards reason in events and engineering reasons in rates. The conversion is three multiplications, and of the three terms only one is genuinely movable.
- Volume is not the lever. Cutting 3000 changes to 400 raises the failure rate per change and lengthens diagnosis, so the product usually gets worse. It is the move boards ask for first.
- Change failure rate is barely the lever. Reaching 2 incidents that way needs something near 0.4%, and published delivery research has consistently placed even strong performers in the tens of percent.
- Escape rate is the lever. Going from 22 to 2 means 24 of every 25 failed changes must be reverted before they are material, which is blast radius and revert time, both architecture.
This also separates it from an error budget. An error budget is availability-denominated and consumed continuously; a change risk budget counts discrete events and attributes them to changes. The two disagree usefully: a 10-minute total outage can blow an incident budget while barely touching an availability one.
Implementation patterns
- Define material before measuring anything: duration, error-rate threshold and which paths count, for example more than 5 minutes above a 1% error rate on a customer-facing path.
- Instrument the escape rate, which almost nobody has. Classify every failed change by whether a customer saw it, for how long and at what error rate. Three months of that replaces the whole estimate, and the assumed 5% is where the error in the model lives.
- Spend on escape-rate mechanisms: canary with automatic rollback on error-rate delta, cell or tenant partitioning so one failure reaches a fraction of users, feature flags with a kill switch independent of the deploy path, and a revert measured in minutes. If material means 5 minutes, a 3-minute automated rollback converts most escapes into non-events by definition.
- Attach a consequence. A team over budget loses the emergency change path, not its deploy rights.
Industry example
Reddit's Pi Day outage on 14 March 2023 ran 314 minutes. Against a 99.99% availability appetite - about 53 minutes a year - that single event is roughly six times the annual allowance; against 99.9%, about 526 minutes a year, it fits with room to spare. An appetite that one plausible incident exceeds is spent by definition and stops guiding anything, which is the first thing to check when a budget is set. The lever there was not approval: what determined impact was that Kubernetes defines no supported downgrade, so the reverse path was a restore from a procedure nobody had exercised.
Failure scenarios
- Material left undefined, so each classification is negotiated by the team that caused the incident.
- Escape rate assumed rather than measured, making the budget a guess with a decimal point.
- The board responds by reducing deploy frequency, and incident count rises.
- Counting events while the real exposure is one tail loss, which the count reassures you about until the event that ends the argument.
Trade-offs
Converting an appetite into rates makes the gap visible, and a visible eleven-times gap is a political problem before it is an engineering one. Measuring escape rate costs classification discipline teams resent. In exchange the organisation stops buying approval and starts buying blast-radius reduction.
When not to use it
Where the real concern is a single catastrophic loss, use a threshold rule per decision; multiplying rates is the wrong model for a tail. Where change volume is a dozen a year, judgement per change is better. And where error budgets are already used, prefer converting the appetite into that currency rather than running two.
Interview question
Q: Your board sets an appetite of two material incidents a year and you ship 3000 changes a year. How do you turn that into something your teams can design against, and what would you refuse to do?
What a strong answer covers: defining material with a duration and an error-rate threshold; the three-term arithmetic and the gap; why volume and change failure rate cannot move far enough; escape rate as the design target with the mechanisms that move it; instrumenting it first because it dominates the error; and refusing to reduce deploy frequency, with the reason.
Quick check
Quiz: 3000 changes a year at a 15% change failure rate and a 5% escape rate - how many material incidents? — About 22, against an appetite of 2. Attack escape rate through blast radius and revert time.
Flashcard: How does a change risk budget differ from an error budget? — An error budget is availability-denominated and consumed continuously; a change risk budget counts discrete material incidents attributed to changes, so a short total outage can blow one and barely touch the other.