A platform has already been through the failure Fastly published in 2021 - a defect shipped on 12 May that a customer's valid configuration armed on 8 June - and has since implemented per-point-of-presence staged activation of customer configurations plus grammar-based fuzzing of the configuration surface. Six months later the same class recurs in a different subsystem. Which gap did that action list never close?
Show the full answer Hide the answer
Separate the two things the action list did
Staged activation bounds the blast radius of the arming event. Fuzzing samples the input space. Neither of them looks at the pairs that actually exist.
Those are different jobs and only the first one was done well. Staged rollout means the next armed defect takes down one point of presence instead of the fleet, which is the highest-value control available and worth every hour spent on it. It does not find the defect; it limits what finding it costs. The team wrote "prevent" on a control that mitigates, and stopped.
Fuzzing looks like the preventive half, and it is much weaker than it appears. The configuration grammar of a programmable platform admits an enormous space, of which the region customers actually author is tiny and highly non-uniform. Random generation spends nearly all of its budget in a region no customer will ever occupy, so it finds parser bugs readily and customer-reachable latent defects rarely.
The gap
Nobody re-tests already-deployed code against configurations written after it shipped. The dangerous object is the pair - this build with that configuration - and the corpus of configurations keeps growing for as long as the build is in production. Fastly's published summary of the 8 June 2021 incident records exactly that shape: the deployment began on 12 May, the triggering configuration arrived 27 days later, and the network detected the resulting errors within a minute. Detection was never the problem, and no pre-release gate can evaluate an input that does not exist yet.
The missing control is continuous verification of the running build against newly authored configurations, evaluated before activation:
- Shadow-evaluate every customer configuration change against the currently deployed build in an isolated fleet, comparing behaviour with the customer's previous configuration. This is a gate on their change against your code, which is the pair nothing else tests.
- Replay the growing corpus against deployed builds on a schedule, so a latent defect surfaces when someone writes the dangerous shape rather than when they activate it.
- Track which (build, configuration) pairs have ever been evaluated together, and publish the size of the unevaluated set. That number is the honest statement of exposure, and most platforms have never computed it.
Why this generalises past edge networks
Read "configuration" as any externally authored input that changes behaviour: tenant settings, rules tables, feature-flag combinations, pricing rules, detection content, uploaded models. Environment parity is normally specified as topology and versions, and the dimension that bites is the input corpus. A staging environment with an identical cluster shape and a dozen hand-written configurations has no parity at all with production's tens of thousands of customer-authored ones. Write down the size and diversity of every input corpus pre-production holds; the difference against production is the list of defect classes it cannot catch.
Common weak answers
- "Extend the canary bake from one day to a week." A bake samples elapsed time. The trigger is an input authored later, so a longer bake would not have helped and it slows every future fix, including the rollback of the next one.
- "Validate configurations more strictly." Rejecting inputs you cannot handle moves the boundary rather than removing it, because the failure is by definition inside the set you believed you handled.
- "More testing before release." The configuration did not exist at release. This is the answer that sounds responsible and changes nothing.
When not to build any of this
If the configuration space is closed and owned by the team that owns the code, configuration is code and the ordinary release process already covers it: enumerate a dozen internal flags in a test and stop. Shadow evaluation, corpus replay and pair tracking are infrastructure, and they pay for themselves only when the input space is open, externally authored and unbounded - which is the defining property of a platform and not of a service.