You have four hours, eight engineers and a payments platform that has never had a game day. Design the exercise. Be specific about what you inject, what you measure, and how you decide whether it succeeded.
Show the full answer Hide the answer
Requirements that drive the structure
Four hours and no prior exercise sets three constraints. The exercise must produce findings even if nothing breaks — otherwise a resilient system yields a wasted afternoon. It must not cause a real incident, because a first game day that does destroys the practice politically. And it must test the humans, not only the software, since on a platform that has never rehearsed, the response process is almost certainly weaker than the architecture.
The design
Two scenarios, ninety minutes each, with a written hypothesis per scenario.
A hypothesis, not a plan to break something: "When the fraud-scoring service returns 5xx for 100% of requests, payment authorisation continues using the cached risk tier, p99 authorisation latency stays under 1.2 s, and the decline rate rises by less than 0.5 points." Three falsifiable predictions. Write them down before injecting anything, because the value of a game day is the gap between what the team predicted and what happened, and that gap cannot be measured retrospectively.
- Scenario 1: a dependency is slow, not down. 2-second added latency on fraud scoring, no errors. Slow is the harder case and the one teams get wrong: breakers trip on error rate, so a slow dependency sails past them while caller threads accumulate.
- Scenario 2: lose a shared dependency the team believes is redundant. One cache node, or one availability zone's worth of a replicated service. The prediction will be "no impact", and this is where hidden coupling surfaces.
Blast radius: one cell, or 5% of traffic by a routing rule, during a low-traffic window. Scope by failure domain rather than by request count — a percentage spread across every cell can still saturate a shared dependency.
Roles, assigned in advance: an incident commander who has not run one before (the exercise is partly for them), a scribe timestamping everything, two engineers injecting, four responding, and one observer who knows the answer and says nothing. The observer silence is what makes it a measurement rather than a lesson.
What to measure
- Time to detect. From injection to the first human noticing. If no alert fired, that is the headline finding.
- Time to correct diagnosis, which is the number that predicts real incident duration.
- The predictions, scored individually. Each one right or wrong.
- What the responders looked at, in order. The scribe's log of this is often the most useful artefact produced, because it shows which dashboard was reached for first and whether it helped.
How you decide whether it succeeded
Success is not "nothing broke". A game day where nothing broke and every prediction held produced one finding: the team's model is accurate, which is worth knowing and is a thin return on four hours.
The exercise succeeded if it produced at least two findings that change something: an alert that did not fire, a runbook step that was wrong, a dashboard that was unreachable, a dependency nobody knew was in the path, or a prediction that was confidently wrong. A failed game day is one where the team learned nothing, not one where the system misbehaved.
Close with 30 minutes of retrospective while it is fresh, every finding written as an owned ticket, and the two or three highest-value ones scheduled before the next exercise. Findings without owners are how this practice dies in its second quarter.
How it fails
- No abort procedure. Define before starting: who calls it, what one action reverts everything, and how long that takes. Test the abort on a no-op injection first.
- Running it during a change freeze or a real incident, which requires checking rather than assuming.
- Injecting five faults to use the time well. Attribution becomes impossible and the exercise produces anecdotes.
- The observer intervening. Understandable, and it converts a measurement into a demonstration.
When not to run one yet
If you have no steady-state signal — no business metric you can read in under a minute to tell whether the platform is healthy — build that first. Without it the exercise cannot distinguish its own injected fault from a real incident starting, and the abort decision becomes a guess. That is a prerequisite, not a parallel workstream.
Equally, if the last three real incidents are still unremediated, run the game day later. A platform with a known unfixed weakness does not need an exercise to find another one, and spending four hours of eight engineers to discover what the postmortems already said is a poor trade.