advanced 4 min answer

A marketplace that runs an annual festive-season sale in the mould of Flipkart's Big Billion Days does chaos engineering twice a year as a staffed game day. Leadership wants it continuous, as a pipeline stage. Sequence the migration, and say where it can go wrong.

chaos engineeringgame dayspipelineresilience testingblast radius
Show the full answer Hide the answer

The sequence

Each step must be reversible, and the ordering is load-bearing: steps 1–3 build the ability to stop, before step 4 introduces anything that needs stopping.

  1. Convert the game-day findings into named hypotheses. A game day produces discoveries; a pipeline stage needs assertions. "When the recommendations service returns 5xx, the product page renders in under 800 ms without recommendations and the error does not reach the user." An experiment without a falsifiable prediction is not a test, it is an outage you scheduled, and this step is what most teams skip.
  2. Build automated steady-state verification first. The pipeline must be able to answer "is the system healthy right now" without a human looking. Concretely: a check against the business metric, not the error rate — orders per minute against the same hour last week. If this does not exist, nothing after this point is safe, and building it is weeks of work that looks like it is not chaos work.
  3. Build the abort. One switch that halts every running experiment and reverts every injected fault, with a measured time-to-stop. Test it by aborting a no-op experiment, repeatedly, until the measured time is known and boring.
  4. Run in the pipeline against pre-production first, with the full automation. This stage finds the automation's bugs rather than the system's, which is exactly what it is for. Expect it to be unrewarding.
  5. Promote one experiment to production, scoped to the smallest blast radius that can still falsify the hypothesis — one cell, one availability zone, 1% of traffic, internal users only. One experiment, not the suite. Run it on a schedule during business hours with people present, and leave it there for several weeks.
  6. Widen the radius before widening the catalogue. Take the single experiment from 1% to 5% to a full cell. Only once one experiment is boring at full scope does a second one join. Teams that add experiments faster than they widen radius end up with a large catalogue of shallow tests that prove nothing.
  7. Gate on the hypothesis, not on the absence of errors. The stage fails when the predicted behaviour does not occur. A fault injection that causes no errors because it was not actually applied must fail the build — a silently ineffective experiment reporting green is the most common way this practice decays into theatre.
  8. Retire experiments. An experiment that has passed every run for a year is testing a property nobody is changing. Move it to quarterly and spend the budget on a new hypothesis.

When this goes wrong

  • Automating before the steady-state check exists. Then the pipeline cannot tell a successful experiment from an outage it caused, and it will cheerfully proceed to the next stage during a real incident.
  • The abort that was never timed. Everyone believes there is a kill switch. Under pressure it turns out to take four minutes, require a deploy, or depend on the control plane the experiment just degraded.
  • Blast radius defined by percentage of traffic rather than by failure domain. 1% of traffic spread across every cell can still take down a shared dependency. Scope by the thing that can break, not by the count of requests.
  • Running during a real incident. The pipeline does not know. It needs a hard interlock against the incident-management system and the change freeze, and a schedule confined to staffed hours.
  • No owner for the findings. A continuous practice generates findings continuously, and without a route into the backlog the same finding is rediscovered every Tuesday until the team mutes the stage.

The point of no return

There is none, which is worth stating plainly: every step is revertible by disabling the stage, and that property should be preserved deliberately. The irreversible thing is organisational — once chaos runs continuously, the team stops doing the twice-yearly game day, and if the automation is then quietly disabled during a busy quarter, the practice has been removed entirely rather than reduced. Keep one staffed game day a year on top, for the scenarios automation cannot cover: a region loss, a failure of the deploy system itself, a human-coordination rehearsal.

How long it really takes

Steps 1–3 dominate and are typically a quarter, most of it the steady-state check. Steps 4–6 are another quarter per handful of experiments. A credible plan is a year to a continuous practice with a modest catalogue, and a plan that promises it next sprint is a plan to skip steps 2 and 3, which is the plan that produces an outage.