advanced 3 min answer

Interview. Your incident runbook says "disable the recommendation service with the feature flag". Walk me through what has to be true for that instruction to work during a real incident, and how you would prove it.

feature-flagskill-switchincident-responsestalenessgame-day
Show the full answer Hide the answer

What the interviewer is testing

Whether you treat a kill switch as a control with its own availability and latency requirements, or as a boolean somebody will remember to flip. The instruction in the runbook is a claim about a distributed system, and most organisations have never tested it.

The clarifying questions that change the answer

  • Where is the flag evaluated? A flag read from a local cache refreshed every 30 seconds behaves completely differently from one evaluated per request against a remote service.
  • How long until the last instance has the new value? The number that matters is the tail, not the mean.
  • What does the flag service do when it is unavailable? Specifically, what a cold-starting instance does when it cannot reach the service during the incident.
  • Does the switch cut the right thing? Disabling recommendations usually means not calling the dependency. If the call is still made and the result discarded, the switch does not reduce load on the thing that is failing.
  • Who can flip it at 03:00, and does that path depend on anything currently broken?

A strong answer's arc

Propagation. With a 30-second poll, a fleet of 400 pods converges over roughly 30 seconds in the good case; with jitter and retries, the tail is a minute or more. During that window the system is in a mixed state, serving two behaviours, and any invariant that assumes all instances agree is violated. If that matters — a pricing rule, a write path — the flag must be evaluated at a single point, or the mixed state must be explicitly safe.

Cold start is where this fails. A pod starting during the incident has no cached value. If the client defaults to "flag off" and off means "call the dependency", the newly started pods ignore your kill switch, and autoscaling guarantees there will be new pods precisely then. The fix is to persist the last known value to local disk and start from it, treating the flag service exactly as a control plane: unavailability freezes configuration rather than reverting it.

Default direction per flag. A kill switch's safe default is enabled-kill, the opposite of a release toggle's. These are different objects with different failure semantics, and storing them in one system with one default is how a flag service outage turns into a product outage.

Evidence. A game day: inject the dependency failure in one cell, flip the switch, and record three numbers — time from flip to last pod converged, the share of requests still reaching the dependency, and whether the business signal recovered. Then kill the flag service and repeat, which is the test that finds the cold-start bug. Findings, not an undisturbed system, are the output.

Common weak answers

  • "We would flip the flag." Describes the intention, not the mechanism, and says nothing about propagation, cold start or default direction.
  • "Evaluate every flag remotely for consistency." Puts the flag service on the critical path of every request, so its availability becomes a hard ceiling on yours, which is the failure the local cache exists to prevent.
  • "We have 300 flags, so we are covered." Flag count is not kill-switch coverage. The relevant question is which dependencies have a tested switch, and the answer is usually a small subset.

What a strong answer adds

Provenance and expiry. A kill switch is a production control, so changes need an audit trail — who, when, previous value — because finance and the postmortem will both ask. And a kill switch that was never exercised decays: the code path behind it rots, the dependency it bypasses becomes load-bearing, and the switch silently stops working. Pair each one with a scheduled exercise and an owner, or accept that the runbook line is aspirational.