advanced 2 min answer

A platform has an error budget policy stating that feature work stops when the budget is exhausted. The budget is exhausted, and product leadership wants a major launch to proceed. How should this be resolved?

error-budgetspolicygovernancereliabilitygitlabscenario
Show the full answer Hide the answer

What the policy is actually for

An error budget converts reliability from an argument into an arithmetic. A 99.9% availability target means roughly 43 minutes of permitted unavailability per month. Spending it is allowed — that is the point. Spending it all means the system has consumed its allowance for risk.

The policy's real function is to make the trade-off explicit and pre-agreed, so the conversation happens when everyone is calm rather than during an incident or a launch week.

Why this situation is not a failure of the policy

It is the policy working. The disagreement is exactly what it was designed to surface. A policy that never creates tension is one that is either never binding or never enforced.

How to resolve it

Not by the reliability team unilaterally blocking the launch, and not by the policy being silently ignored. Both outcomes destroy the mechanism — the first makes reliability an obstruction to be routed around, the second makes the policy theatre.

The resolution is an explicit, documented exception decided by whoever owns the business outcome:

  1. Present the actual risk, not the policy. What is the current failure rate, what is causing it, what is the probability the launch makes it worse, and what is the expected customer impact.
  2. Name the decision-maker. Someone accountable for both reliability and revenue — typically a director or VP, not the engineering team and not the product manager alone.
  3. Attach conditions. If the launch proceeds: a reduced blast radius, a slower rollout, an explicit rollback trigger, and a commitment to reliability work immediately afterwards with a named owner and a date.
  4. Record it. The exception, the rationale and the conditions, so a pattern of exceptions is visible.

The signal to watch

One exception is a business decision. A pattern of exceptions means the SLO is wrong.

If the budget is exhausted every month and the launch always proceeds, then the organisation has revealed that it values feature velocity above the stated target. The honest response is to lower the SLO to something the organisation will actually defend, or to invest in reliability so the budget stops being exhausted. Continuing to declare a target nobody enforces corrodes every other commitment.

The organisational point

Error budgets fail when they are owned by the reliability team as a veto. They work when they are owned jointly, with an agreed exception path and visible history — because their value is not the enforcement, it is the shared vocabulary for a trade that was previously made implicitly by whoever pushed hardest.