Chaos as a Test
Fault injection with a hypothesis, a blast radius and an abort condition.
4 to work through
-
intermediate
Leadership wants to start chaos engineering. What is your first experiment, and what must be true before you run anything?
2 min answer -
advanced
How does chaos engineering differ from a resilience test, and how should it be integrated into delivery?
2 min answer -
advanced
Walk me through this one. You run the resilience engineering team at a global edge network in the mould of Akamai. Two years in you have 400 automated fault-injection experiments running weekly, no customer-visible incident has ever been caused by one, and the last genuinely new finding was six months ago. Your VP asks whether to keep funding the team. What do you say?
3 min answer -
advanced
Your SRE team wants to run fault injection in production. Leadership is nervous. How do you make the case and what do you insist on?
2 min answer
3 terms in this topic
Chaos Experiment
A controlled fault injected into a system to test a written hypothesis about how it should degrade.
metricChaos Experiment Yield
New findings per unit of resilience-engineering effort - the number that separates a chaos programme's decaying discovery value from its cheap and pe…
conceptSteady-State Hypothesis
The measurable statement of normal behaviour that a chaos experiment predicts will hold while a fault is injected, without which the exercise is not …
Neighbouring topics
Testing & Quality Architecture
General material on designing a testing strategy as an architectural concern.
Test Architecture Strategy
Choosing what to verify where, given the failure modes that actually occur.
Test Pyramid Shapes
Pyramid, trophy and honeycomb, and the system properties that justify each shape.
Integration Test Boundaries
What sits inside a test's boundary, what is faked, and the confidence that follows.
Contract Testing at Scale
Keeping dozens of services compatible without an environment that runs all of them.
Consumer-Driven Contracts
Consumers declaring what they rely on, and providers verifying against those declarations.
Test Data Management
Realistic data without copying production personal data into a weaker environment.
Synthetic Data
Generating data with the shape and edge cases of the real thing, and where it misleads.
Environment Parity
The differences between staging and production that decide which bugs survive to release.
Service Virtualisation
Standing in for a dependency you cannot call, and keeping the stand-in honest.
End-to-End Test Economics
Why broad end-to-end suites get slow, flaky and abandoned, and what to keep.
Non-Functional Test Strategy
Testing availability, latency, security and recovery rather than only behaviour.
Performance Test Design
Workload models, warm-up, think time, and the distribution the average hides.
Security Testing in the Pipeline
SAST, DAST, dependency and secret scanning, and what to do with the findings.
Accessibility Testing
Automated checks, their ceiling, and the manual testing that has to sit above it.
Mutation Testing
Measuring whether tests would actually notice a defect, not just cover a line.
Flaky Test Management
Quarantine, detection, and the trust a suite loses once red stops meaning broken.
Testing in Production
Synthetic transactions, dark launches and shadow traffic, done deliberately and safely.
Quality Gates
Thresholds that block a release, who may override them, and how they decay.