SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
4 to work through
-
intermediate Multiple choice
A brokerage sets a single availability target of 99.9% for its whole platform. Why is that wrong, and what should replace it?
2 min answer -
advanced
A live-streaming platform defines an SLO of "99.9% of API requests succeed". During a major event, the SLO is met while viewers report the product is unusable. What is wrong with the SLO?
2 min answer -
advanced
A service exhausts its error budget in week two of a 28-day window. What actually happens next?
2 min answer -
advanced
Leadership asks for "five nines" across the platform. Engineering says it is impossible. Design the response.
2 min answer
5 terms in this topic
Aspirational and Achievable SLO
The distinction between the reliability a team wishes for and the reliability its current architecture and dependencies can actually deliver.
practiceJourney-Level SLO
Setting reliability targets per user journey rather than per platform or per service, because consequence differs by journey and only a journey-level…
metricRequest-Based and Window-Based SLI
Two ways of computing the same reliability target — counting good events, or counting good time windows — which produce materially different numbers.
metricSLI, SLO and SLA
The measurement, the internal target, and the external contract — with the error budget as the mechanism that makes the target consequential.
practiceUser Journey SLO
Defining reliability targets around what a user is trying to accomplish, measured where the user is, rather than around per-request success rates at …
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.