Walk me through this one. You run the resilience engineering team at a global edge network in the mould of Akamai. Two years in you have 400 automated fault-injection experiments running weekly, no customer-visible incident has ever been caused by one, and the last genuinely new finding was six months ago. Your VP asks whether to keep funding the team. What do you say?
Show the full answer Hide the answer
What the interviewer is testing
Whether you can separate a chaos programme's two products - regression assurance and discovery - and put a number on each. Candidates who defend the programme as one indivisible good, or who concede it should be cut because it found nothing, both fail. They also want to see whether you treat "no incidents caused" as evidence of value, which it is not: it is evidence the safety controls work.
The clarifying questions that change the answer
- What is the organisation's change rate? Discovery yield tracks architectural churn. A platform mid-migration has a growing hypothesis space; a stable one does not.
- Are the 400 experiments spread across failure domains or clustered? Four hundred variations of "kill an instance" is one experiment run four hundred ways.
- What has happened to time-to-detect over two years? That is the property fault injection measures that nothing else does, and if it improved, the programme has a result independent of findings.
- Which failure classes have never been injected? Before concluding the space is mined, check it against the list nobody gets to: a dependency returning wrong but well-formed data, a dependency that is slow rather than down, clock skew, an expired certificate, a partially applied deployment, and the control plane rather than the data plane.
A strong answer's arc
Split the budget in two, because the two products have different cost curves.
Regression. The 400 experiments are now cheap: they consume machine time and an alerting integration, not headcount. They are a plausible reason the last six months were quiet, and deleting them re-opens every failure mode they pin down. Move them into the platform team's normal pipeline and keep them. Cost: small and flat.
Discovery. This is the part that has stalled, and the honest reason is that the hypothesis space of a stable architecture is finite and you have mined the accessible region. Measure it as findings per engineer-week and show the decay curve. If it has been near zero for six months at full staffing, the current allocation is wrong.
The recommendation: shrink discovery to a small standing capability, and re-fund it on architectural events - a new region, a datastore migration, a new dependency, a cell split - because those are the moments the space reopens and the moments an outage is most likely. Pair each such event with a required set of experiments as part of its readiness, which makes resilience work a gate on change rather than a standing team looking for work.
Common weak answers
- "Chaos engineering is always worth it." Every practice has a yield curve. Refusing to measure yours is how budgets get cut by someone who did measure it.
- "Zero incidents proves it works." You cannot attribute an absence. The defensible claims are the specific failure modes the experiments demonstrated and the time-to-detect improvement.
- "Run more in production." Blast radius is not the constraint here; the supply of hypotheses is.
What a strong answer adds
Name the second deliverable the programme has been producing without billing for it: runbook accuracy and on-call decision speed, measurable as time-to-decide in game days. And say where you would stop entirely - an organisation deploying monthly, in one region, with untested backups, gets more from a restore drill than from any fault injection, and recommending that against your own budget is the answer that gets believed.