Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
4 to work through
-
advanced
A brief database slowdown caused a two-hour full outage. Explain the likely amplification chain and the fixes at each stage.
2 min answer -
advanced
A travel platform depends on hundreds of external suppliers with varying reliability. How should resilience testing be designed when the failures originate outside the system?
2 min answer -
advanced
How would you test that your service degrades correctly when a dependency's latency rises tenfold?
2 min answer -
advanced
Review this resilience programme. Chaos experiments run weekly in staging at 03:00, they inject only instance termination, results are recorded in a spreadsheet, and there is a kill switch that has never been used. What would you change, and what would you keep?
3 min answer
2 terms in this topic
Failure Injection Testing
Deliberately introducing faults into a system under test to verify that timeouts, retries, fallbacks and circuit breakers behave as designed.
practiceResilience Testing
Verifying that failure-handling behaves as designed — a category of testing distinct from functional and load testing, and usually absent.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.