On 22 July 2021 a software configuration update triggered a bug in Akamai's Edge DNS, and services were disrupted for roughly an hour until the update was rolled back. Consider a system whose control plane — the thing that distributes configuration — becomes unavailable while the data plane is untouched. What should happen to traffic?
Show the full answer Hide the answer
Second by second, what happens
With static stability (the correct design). The control plane goes away. Every data-plane node is already holding a complete local copy of the configuration it needs, so it keeps answering. Traffic is unaffected. The only thing lost is the ability to change anything — no new routes, no new certificates, no scaling decisions. Nodes log that their configuration is stale and emit a staleness metric. The incident is a change freeze, not an outage.
Without it. Nodes cannot confirm their configuration is current, so they stop serving, or they fall back to a default that does not match reality, or they queue while retrying. The retries arrive at a control plane that is already failing, and the data plane's load becomes the reason the control plane cannot recover. A configuration-distribution problem has become a total outage, and the outage now actively prevents its own repair.
Where it amplifies
The dangerous amplification is the retry loop above: a dependency on the control plane for serving means that control-plane trouble generates data-plane traffic toward the control plane at exactly the wrong moment. This is why "retry and queue" is not merely slower but actively harmful.
The second amplification is scope. A control plane is shared by construction — that is what makes it a control plane — so its failure is correlated across everything. A design in which control-plane health gates data-plane serving has quietly made every node's availability depend on one component's.
What the user sees
With static stability: nothing, for the duration. Without it: a total outage on a system whose serving path never broke.
Why the other options fail
- "Fall back to a safe default that denies unknown routes." Fail-safe reasoning applied in the wrong place. Deny-by-default is right for authorisation; for configuration it means discarding correct state you already have in favour of less information. The last known good configuration was working a second ago. There is no safety argument for preferring a default to it.
- "Retry and queue until configuration is confirmed fresh." The intuitive engineering answer and the one that converts an hour of frozen changes into an hour of downtime, while adding load that delays recovery. Freshness of configuration is not a correctness requirement for the overwhelming majority of requests.
- "Shut down because it cannot verify currency." The strictest reading of correctness and almost always wrong. It treats a stale route as more dangerous than no service, which is true for a small set of cases — a revoked credential, a legally-mandated block — and false for routing, load balancing and discovery. Handle those few cases with short-TTL deny lists carried separately, rather than by making everything fail closed.
What stops it
- Every data-plane node holds a complete local copy of what it needs, on local disk, loaded at startup from cache if the control plane is unreachable. Starting cold without the control plane is the case to test, because "it keeps running" is easy and "it can restart" is where designs fail.
- Pre-provision capacity rather than scaling reactively. Static stability includes not needing the control plane to add capacity: run with enough headroom that a zone loss needs no new instances. That is the expensive half of the idea and the reason it is a trade-off rather than a free win.
- Alert on staleness as a distinct signal, with a threshold measured in hours. Stale and serving is a tolerable state that must be visible, or a configuration freeze will go unnoticed for a week.
- Separate the config push's blast radius. Akamai's incident was triggered by a configuration update; the companion control is staged rollout of configuration itself, so a bad push reaches a fraction before it reaches everything.
When not to do this
Staleness has a real cost, and for some state it is unacceptable. A revoked certificate, a cancelled authorisation, a blocked account or a GDPR erasure must propagate, and "keep serving the last known good" is the wrong answer for them. The resolution is not to abandon static stability but to split the configuration by how much staleness it tolerates: routing and discovery serve indefinitely from cache, while a small deny list carries a short TTL and fails closed when it expires.
The other case is cost. Static stability requires the headroom to absorb a failure without provisioning, which means paying for capacity you do not use — commonly in the region of 30–50% in an N+1 regional design. For a service where a few minutes of unavailability is genuinely acceptable, reactive scaling is cheaper and the honest choice.