Reliability & Resilience
General material on designing for failure.
6 to work through
-
intermediate
A team is launching a new service and asks what its SLO should be. How do you help them decide, and why is "99.99%" usually the wrong first answer?
2 min answer -
intermediate
The business asks for "100% uptime" for a new customer portal. Walk me through the conversation that ends in an agreed SLO.
3 min answer -
advanced
A super-app combines messaging, payments, social feeds, mini-programs and notifications in one product used by hundreds of millions daily. What is the dominant reliability risk, and what structural property addresses it?
2 min answer -
advanced
A workflow executes over days or weeks while individual workers restart many times. How should workflow state, timers, retries, heartbeats and task ownership be designed?
2 min answer -
advanced
In July 2019 a single regex deployed globally took Cloudflare's network to 100% CPU within seconds. What does this say about how configuration should be released?
2 min answer -
advanced
Maersk rebuilt roughly 4,000 servers and 45,000 PCs in about ten days after NotPetya in 2017, and recovered its directory only because one data centre had been offline during the attack. What does this say about DR design?
2 min answer
14 terms in this topic
Availability Calculation
Deriving a system's availability from its components, remembering that dependencies in series multiply.
practiceBlameless Postmortem
An incident review that seeks the systemic conditions that made a failure possible, explicitly excluding individual fault.
practiceCapacity Planning
Deciding in advance how much capacity will be needed, given growth, seasonality and failure scenarios, and ensuring it can be there in time.
practiceChaos Engineering
Deliberately injecting failure into a system to discover, before an incident does, which of your resilience assumptions are false.
practiceDisaster Recovery
The plan and capability for restoring service after an event that takes out a whole site, region or system.
patternDurable Execution
Persisting a workflow's progress outside the process executing it, so that worker restarts, deployments and crashes resume from the last completed st…
metricError Budget
The amount of unreliability an SLO permits, treated as a resource that feature velocity spends.
practiceGame Day
A scheduled exercise in which a failure is deliberately introduced and the team responds as though it were real, to test the system and the response …
practiceIncident Command
Assigning explicit roles during an incident — commander, operations lead, communications lead, scribe — so coordination does not compete with diagnosis.
metricRTO and RPO
How long recovery may take (RTO) and how much data may be lost (RPO), the two numbers that determine the cost of a resilience design.
conceptService Level Agreement
A contractual commitment about service level, with a defined remedy — usually a service credit — when it is missed.
metricService Level Indicator
The actual measurement of a service's behaviour that an objective is set against — a ratio of good events to valid events.
metricService Level Objective
An internal target for a service level indicator, set below the level at which users notice, and used to decide whether to ship or to stabilise.
practiceSeverity Levels
A predefined scale of incident impact that determines who is woken, how fast, and what process applies.
Neighbouring topics
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.