Distributed Systems 06 Oct 2026 27 min read

Serving without the control plane

How production data planes are designed to keep serving while the coordinator that instructs them (consensus store, registry, xDS server, failover manager) is down, and the one situation in which continuing to serve is the wrong move.

Kubernetes, Envoy, Eureka and Patroni all answer the same question in their repositories: what may the data plane do with its last instructions when the control plane goes silent? This guide reconstructs the four recurring mechanisms (last-known-good caches, disbelief thresholds, leases with fences, failure-time-only coordination), prices them with shipped defaults, and shows from GitLab's September 2026 incident what happens when fail-static meets an exclusive role without a fence. A reader leaves able to classify every config their data plane holds as safe-stale or dangerous-stale and to defend a staleness budget in review.

The finding that surprised me

Envoy, Kubernetes and Eureka each ship a numeric threshold (50%, 55%, 15%) above which the system stops believing its own health signal, three independent implementations of one unnamed rule: a signal reporting mass death is more likely reporting its own failure.

What you get out of it

  • Fail-static is the industry default and is written down as a Kubernetes founding principle: components continue to do what they were last told in the absence of new instructions.
  • Three systems independently converge on a disbelief threshold (Envoy 50%, Kubernetes 55%, Eureka 15%) above which mass-death health signals are ignored rather than acted on.
  • The default flips with exclusivity: Patroni demotes a primary the moment the lock update fails, and its failsafe mode requires unanimous member acknowledgement because DCS quorum and member quorum are different views.
  • A lease bounds the decision, not the behaviour: GitLab's primary kept committing for 33 minutes after losing a 30-second lock because fencing was configuration rather than a kernel watchdog.
  • Coordination demand spikes exactly during failures (Physalia calls EBS a sometimes-coordinating system), so a coordinator sized for normal operation browns out on the day it matters, as EBS did in 2011.

Scope

Why this, now. GitLab's INC-14518 corrective actions, published in the weeks before this guide, are the freshest public record of the half-dead-leader failure class, and the September 2026 publication of concrete numbers (a 33-minute unfenced divergence window against a 30-second lease) makes the staleness-budget argument unusually concrete.

What it does not cover. The change plane (bad instructions propagated by deploys and config rollouts), client-side lock correctness and fencing tokens (covered by the earlier when-two-hold-the-lock guide), consensus protocol internals, and everything published only on unreachable hosts: this session's network reached code hosts only, so the AWS Builders' Library static-stability essays, the Roblox 2021 Consul outage and the Cloudflare 2020 etcd outage appear only where reachable artefacts restate their mechanisms, and there are no conference talks.

Open the field guide → Self-contained: it loads nothing at read time, follows your system theme, and prints cleanly.