Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
5 to work through
-
intermediate
A mobility platform wants to start chaos engineering. Its dispatch system is business-critical and the team is nervous. What is a responsible starting sequence, and what must exist first?
2 min answer -
intermediate
A team wants to adopt chaos engineering after reading about it. What must be in place first, what is the correct progression, and what does chaos engineering not tell you?
3 min answer -
advanced
A dependency fails only under high load, so normal testing never exposes it. How should load testing, fault injection, capacity experiments and traffic shadowing be combined to find it?
2 min answer -
advanced
Leadership read about Chaos Monkey and wants chaos engineering in production next month. Your last three incidents were caused by known unfixed reliability issues. What do you say?
2 min answer -
advanced
What is the first chaos experiment you would run on a system you have just inherited, and what must be in place first?
2 min answer
3 terms in this topic
Chaos Engineering in Practice
Deliberately injecting failure into production to verify that resilience mechanisms work — an experiment with a hypothesis, not vandalism.
practiceLoad-Coupled Fault Injection
Injecting dependency faults while the system is at realistic peak load, because timeouts, pool sizes and breaker thresholds only reveal whether they …
case-studyNetflix: Chaos Monkey and the Simian Army
Netflix deliberately terminated production instances during business hours to force engineers to build for failure rather than hope against it.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.