advanced 3 min answer

Fastly began a software deployment on 12 May 2021 that ran for nearly a month without incident. On 8 June a customer pushed a valid configuration change and 85% of the network started returning errors. What class of risk does that expose, and what does it mean for deployment safety gates?

fastlyedgeconfigurationlatent-defectblast-radiuspostmortem
Show the full answer Hide the answer

The situation they were in

Fastly's own account of the 8 June 2021 outage is specific: a deployment begun on 12 May introduced a bug that could be triggered by a particular customer configuration under particular circumstances. On 8 June a customer pushed a valid configuration change that met those circumstances, and 85% of the network returned errors. Fastly detected the disruption within one minute, isolated the cause and disabled the configuration, and its engineering leadership stated that 95% of the network was operating normally within 49 minutes.

Why a staged deployment would not have caught it

The deployment was staged, observed and uneventful. The defect's activation was decoupled from its release by 27 days, because the trigger was an input, not a code path exercised at rollout. That breaks the signal most release processes rest on: time since deploy with no errors. Here that signal was four weeks of false reassurance.

An edge fleet has two control planes changing at different rates. Code ships on the platform's schedule, under canaries. Configuration ships on thousands of customers' schedules, continuously, and is trusted because it is valid. The second plane is where the trigger lives, and it usually has no canary at all.

What this changes about release gates

  • Canary the configuration, not only the code. A config change rolled to a small slice of the fleet first, with an automatic revert on error rate, turns a global event into a local one.
  • Replay the real configuration corpus against a new build before it goes global. The corpus is the input space, and it is the one input space a platform actually owns a copy of.
  • Keep a disable path that does not traverse the broken component. Fastly's recovery was fast because the triggering configuration could be switched off independently. A kill switch that needs the failing data path is not a kill switch.
  • Alert on the activation, not the deployment. The question "what changed?" must search customer configuration changes, not only internal releases, or the investigation starts in the wrong place.

What it cost them

An hour of global errors across a large share of the web, followed by the slower tail that any edge fleet pays after mass errors: caches come back cold, so origins absorb a load spike after the fix, and customer-visible latency recovers more slowly than the error rate does.

Common weak answers

  • "They should have tested more." The build passed its tests and ran for 27 days. Name the missing test - replaying real customer configurations against a candidate build - or the answer is a sentiment.
  • "Roll out configuration more slowly." Customer configuration changes are customer-visible latency; a platform that takes an hour to apply a config has sold a different product. The staging has to be fast and automatic, not manual.
  • "Add more monitoring." Detection was already one minute. The variable worth moving was blast radius.

Where copying the response would be a mistake

One-minute detection and a 49-minute recovery rest on a uniform global fleet, one configuration plane and the authority to disable a paying customer's configuration without a meeting. A platform with twenty services and a change-approval queue cannot buy that, and chasing the detection number is the wrong lesson. The transferable part is cheaper: slow the rollout of the plane you control, canary the plane you do not, and keep one tested switch that turns a trigger off. For a small platform, a slower rollout beats faster detection, because the blast radius was the variable that mattered.