Error Budgets
Unreliability as a resource that feature velocity spends.
4 to work through
-
advanced
A platform has an error budget policy stating that feature work stops when the budget is exhausted. The budget is exhausted, and product leadership wants a major launch to proceed. How should this be resolved?
2 min answer -
advanced
A team consistently ends every window with 90% of its error budget unspent. What does that tell you?
2 min answer -
advanced
A team consistently exhausts its error budget and continues shipping features. What has gone wrong, and what would make the budget actually function?
2 min answer -
advanced
How does an error budget actually change behaviour, what makes it fail in practice, and what should happen when it is exhausted?
3 min answer
3 terms in this topic
Error Budget
The permitted amount of unreliability implied by an SLO, used as an automatic arbiter between shipping features and improving reliability - and, when…
practiceError Budget Policy
The pre-agreed consequences of exhausting an error budget, which is what turns an SLO from a metric into a decision-making mechanism.
case-studyGoogle SRE: Error Budgets as a Negotiation Device
Google resolved the standing conflict between shipping speed and reliability by giving both sides a shared number and a pre-agreed consequence.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.