practice

Chaos Experiment

A controlled fault injected into a system to test a written hypothesis about how it should degrade.

chaosresiliencehypothesis

The discipline that distinguishes chaos engineering from breaking things is that it is a hypothesis test with a stated expectation and an abort condition, run in a bounded blast radius. "We believe that if the recommendations service becomes unavailable, the product page still renders in under 500 ms with a static fallback and no error rate increase on checkout."

Running it is where value appears, because the results are so often surprising: the fallback exists but was never wired up, the timeout is 30 seconds rather than 300 ms, a dependency thought to be optional is on the critical path, or the failure is handled correctly and nobody is alerted, which is its own finding.

The sequence that keeps it safe: hypothesis, smallest environment that can falsify it, defined abort criteria, an announced window early on, then progressively larger scope and eventually production with real blast-radius control.

Two prerequisites that teams skip. Do not run experiments on a system with known unaddressed reliability problems — you already know it breaks, and the finding buys nothing. And ensure observability is good enough to see the effect, since an experiment you cannot measure is just an outage you caused.