Serving without the control plane
How production data planes are designed to keep serving while the coordinator that instructs them (consensus store, registry, xDS server, failover manager) is down, and the one situation in which continuing to serve is the wrong move.
Kubernetes, Envoy, Eureka and Patroni all answer the same question in their repositories: what may the data plane do with its last instructions when the control plane goes silent? This guide reconstructs the four recurring mechanisms (last-known-good caches, disbelief thresholds, leases with fences, failure-time-only coordination), prices them with shipped defaults, and shows from GitLab's September 2026 incident what happens when fail-static meets an exclusive role without a fence. A reader leaves able to classify every config their data plane holds as safe-stale or dangerous-stale and to defend a staleness budget in review.
Envoy, Kubernetes and Eureka each ship a numeric threshold (50%, 55%, 15%) above which the system stops believing its own health signal, three independent implementations of one unnamed rule: a signal reporting mass death is more likely reporting its own failure.
What you get out of it
- Fail-static is the industry default and is written down as a Kubernetes founding principle: components continue to do what they were last told in the absence of new instructions.
- Three systems independently converge on a disbelief threshold (Envoy 50%, Kubernetes 55%, Eureka 15%) above which mass-death health signals are ignored rather than acted on.
- The default flips with exclusivity: Patroni demotes a primary the moment the lock update fails, and its failsafe mode requires unanimous member acknowledgement because DCS quorum and member quorum are different views.
- A lease bounds the decision, not the behaviour: GitLab's primary kept committing for 33 minutes after losing a 30-second lock because fencing was configuration rather than a kernel watchdog.
- Coordination demand spikes exactly during failures (Physalia calls EBS a sometimes-coordinating system), so a coordinator sized for normal operation browns out on the day it matters, as EBS did in 2011.
Scope
Why this, now. GitLab's INC-14518 corrective actions, published in the weeks before this guide, are the freshest public record of the half-dead-leader failure class, and the September 2026 publication of concrete numbers (a 33-minute unfenced divergence window against a 30-second lease) makes the staleness-budget argument unusually concrete.
What it does not cover. The change plane (bad instructions propagated by deploys and config rollouts), client-side lock correctness and fencing tokens (covered by the earlier when-two-hold-the-lock guide), consensus protocol internals, and everything published only on unreachable hosts: this session's network reached code hosts only, so the AWS Builders' Library static-stability essays, the Roblox 2021 Consul outage and the Cloudflare 2020 etcd outage appear only where reachable artefacts restate their mechanisms, and there are no conference talks.
Other field guides
Your write hasn't happened here yet
Reconstructs the four-part machine behind read-your-writes (mint a position token at commit, carry it with the client, route to a caught-up replica, …
24 sources · 16 organisations · 4 postmortemsHow long to wait: timeouts and deadlines in distributed systems
A field guide to timeouts and deadlines: the difference between a local per-call timeout and a propagated deadline, the stack of independent timers e…
22 sources · 13 organisations · 3 postmortemsWhen two hold the lock: what the repositories admit about distributed mutual exclusion
Every mainstream distributed lock ships with a written admission that it cannot guarantee mutual exclusion, and this guide reads those admissions whe…
22 sources · 13 organisations · 3 postmortems