Chaos Experiment Yield
also called Findings per Engineer-Week, Resilience Discovery Rate
New findings per unit of resilience-engineering effort - the number that separates a chaos programme's decaying discovery value from its cheap and permanent regression value, and decides which half to fund.
Two years into a chaos programme, a team has 400 automated experiments running weekly, no customer-visible incident ever caused by one, and no new finding in six months. Someone asks whether it is still worth the money, and the debate immediately goes badly, because both sides are arguing about one thing that is actually two.
A chaos programme produces regression assurance and discovery, and they have opposite cost curves. Regression is the 400 experiments, now consuming machine time rather than headcount, pinning failure modes that would otherwise silently return. Discovery is the generation of new hypotheses about how the system might fail, and it is people-expensive.
Experiment yield measures only the second one: new findings per engineer-week. It falls over time in a healthy programme, and reading that fall as failure is how good programmes get cut.
Why it matters
Without the split, the two available arguments are both wrong. "We have caused no incidents, so it is working" attributes an absence, which is not evidence of anything except that the safety controls hold. "We have found nothing in six months, so stop" deletes the regression suite along with the stalled discovery effort, and re-opens every failure mode the 400 experiments hold down.
The hypothesis space of a stable architecture is finite. Yield decays because the accessible region has been mined, not because the practice stopped working, and it reopens on architectural change. That makes yield a scheduling signal rather than a verdict.
Implementation patterns
- Count findings, and define one. A finding is a hypothesis falsified: the system did not degrade as predicted. A passed experiment is not a finding, and neither is a flaky test discovered along the way.
- Report yield against effort, not against experiment count. Four hundred variations of "kill an instance" is one experiment run four hundred times, and counting them flatters the programme.
- Track time-to-detect separately. It is the property fault injection measures that nothing else does, it improves even when findings do not, and it sits in front of every other recovery metric - a fault spotted in 4 minutes and one spotted in 50 minutes are the same technical result and a different outage.
- Convert proven experiments into scheduled regression owned by the platform pipeline, so their cost becomes infrastructure rather than headcount and they stop competing with discovery for people.
- Re-fund discovery on architectural events - a new region, a datastore migration, a new critical dependency, a cell split - and attach a required experiment set to each as a readiness gate.
- Audit the unexplored classes before concluding the space is mined: a dependency returning well-formed but wrong data · slow rather than down · clock skew · an expired certificate · a partially applied deployment · the control plane rather than the data plane. Most programmes that report zero yield have injected none of these.
Industry example
The shape is characteristic of large edge and delivery networks - an organisation in the mould of Akamai, running many points of presence with a stable data-plane architecture and a high rate of configuration change. Yield in the first year is high because the accessible failure modes are plentiful; by year two the data plane's hypotheses are largely exhausted while the control plane and the configuration pipeline are barely touched, and the programme's measured yield collapses even though a large region of the space is untested. The corrective is an explicit coverage map of failure classes against system components, reviewed alongside the yield number.
Failure scenarios
- The programme is cut on a low yield number and the regression experiments go with it, so failure modes quietly return.
- Yield is gamed by counting re-runs, minor variants or unrelated bug discoveries as findings.
- Discovery continues at full staffing against a mined space, producing activity and no result, which eventually gets noticed by someone with a budget spreadsheet rather than by the team.
- Nobody measures time-to-detect, so the programme's most defensible result is invisible.
- Experiments only ever injected into the data plane, so a control-plane outage is the first of its kind in production.
Trade-offs
Measuring yield invites a bad manager to cut a programme on a single number, and that risk is real. The alternative - refusing to measure - means the cut happens anyway, decided by someone who did the arithmetic without you. The defensible position is to publish both lines, the flat cheap regression cost and the decaying discovery cost, with a coverage map that shows what remains unexplored and a stated trigger for re-investment.
When not to use it
A programme in its first year should not be measured this way: yield is high and noisy, and the metric's purpose is to detect decay that has not happened yet. More importantly, an organisation deploying monthly, in one region, with backups that have never been restored, should not have a chaos programme to measure. A restore drill and a dependency map return more per hour than any fault injection, and recommending that against your own budget is the analysis that gets believed.
Interview question
Q: Your VP wants to cut the resilience engineering team because it has found nothing in six months. You think part of that is right. Make the argument.
What a strong answer covers: split regression from discovery and cost them separately · keep the 400 experiments because they are now machine cost and hold down known failure modes · show the discovery decay curve as findings per engineer-week and explain why a finite hypothesis space produces it · audit the never-injected classes before conceding the space is mined · propose shrinking discovery to a standing capability re-funded on architectural events · and name time-to-detect and runbook accuracy as the deliverables the programme has been producing without billing for them.
Quick check
Quiz: Why is "we have caused no customer-visible incidents" not an argument for a chaos programme's value? It attributes an absence; it is evidence the blast-radius controls work, not that the experiments found anything.
Flashcard: Why does chaos experiment yield fall in a healthy programme? - The hypothesis space of a stable architecture is finite and the accessible region gets mined; yield reopens on architectural change, so it is a scheduling signal rather than a verdict.