Fastly's network pushes customer configuration globally in seconds, and that speed is part of the product. On 8 June 2021 a valid customer configuration change triggered a bug introduced by a deployment that began on 12 May, and about 85% of the network returned errors. After the incident someone proposes staging every configuration push in waves with a 10-minute bake. What does that buy and what does it pay?
Show the full answer Hide the answer
What is gained
Containment, and only if the waves are drawn around infrastructure rather than customers. A push that reaches 5% of points of presence first turns "85% of the network returns errors" into "one wave returns errors", which is the difference between a global outage and a regional one.
Note what the 10-minute bake does not buy here. Fastly's own account of 8 June 2021 has the disruption beginning at 09:47 UTC, detection within about a minute, the triggering configuration identified and disabled by 10:27, and 95% of the network healthy within 49 minutes. For a failure that announces itself in one minute, fast reversal already provides what a bake time is supposed to provide.
What it costs
- The product. Customers use fast propagation to fix their own incidents: switching a failing origin, adding a block rule, correcting a redirect. A 10-minute staged push converts a 10-second mitigation into a 10-minute one for every customer, including the ones who are mid-incident.
- Config skew as a new failure class. During the bake, one customer's traffic is served by locations running two different configurations. Anything that assumes a single consistent view breaks in a new way: purges that half-apply, signed-URL key rotation, cookie-based routing, A/B assignment.
- Operational cost. Staged propagation needs wave definitions, health evaluation per wave, and an abort path, for a change type that pushes orders of magnitude more often than code.
The distinction that actually matters
The 2021 trigger was a customer configuration that was customer-scoped in intent and platform-scoped in effect, because it was interpreted by shared code running on every node. The code had been in production for 27 days without anyone exercising that path.
That splits configuration into two classes with different speed limits:
- Parameterising config — a new origin host, a TTL, an ACL entry — reaches only code paths that already run under live traffic for thousands of customers. Keep the fast path; the blast radius is the customer's own traffic.
- Path-activating config — the first use of a newly shipped feature, a new construct in the config language, a flag whose branch has never executed in production. Stage this, because the risk is the latent code and not the configuration.
Classifying automatically means the config compiler must know which features a change reaches and which of those have never been executed with production traffic. That is real engineering, and it is the thing worth building rather than a blanket delay.
The design I would propose
Keep seconds-level propagation as the default. Require staged propagation when the compiled configuration reaches a feature whose first-execution count in production is zero, which also makes the latent-code problem measurable instead of anecdotal. Keep the disable path for a single customer's configuration fast and one-click, since that is what ended the 2021 incident.
When this is the wrong answer
When the failure class is silent or delayed, invert the choice. A wrong rounding rule, a slow memory leak, a cache poisoned with stale entries: none of these show up in a minute, so fast reversal is worth nothing because nobody knows to reverse. Those changes need bake time measured against how long the fault takes to appear, and accepting slower propagation is the correct price. Choose between propagating slowly and reversing fast by how loud the failure class is, not by how severe it would be.