practice

Error Budget Policy

The pre-agreed consequences of exhausting an error budget, which is what turns an SLO from a metric into a decision-making mechanism.

sloreliabilitygovernance

An SLO without a policy is a dashboard. The policy is the part that changes behaviour, and it must be agreed before the budget is exhausted, when the discussion is hypothetical and reasonable rather than under pressure with a launch scheduled.

A typical policy escalates. Budget healthy: ship freely, and consider whether reliability is being over-invested in. Budget half consumed: review and prioritise reliability work alongside features. Budget exhausted: feature releases pause, engineering effort moves to reliability until the budget recovers, and any exception is a named decision at a defined level.

What this achieves is the resolution of the standing tension between velocity and stability without either side needing to win an argument each time. Reliability work becomes automatically prioritised exactly when it is needed, and — the half that is usually forgotten — a consistently unspent budget is evidence that the target is too conservative and that the organisation is buying reliability nobody asked for.

The precondition is that the SLO measures something users actually experience. Freezing feature work because a badly chosen indicator dipped will destroy the mechanism's credibility in one cycle.

The organisational precondition is genuine agreement from product leadership. A policy that is overridden the first time it binds is worse than none, because it has established that the mechanism is decorative.