Serving without the control plane  / field guide
Practitioner field guide · 2026-10-06

Serving without the control plane

Every data plane takes instructions from a smaller system that can die: a consensus store, a service registry, an xDS server, a failover manager. This guide reconstructs, from the repositories of Kubernetes, Envoy, Eureka, Patroni and etcd, from GitLab’s September 2026 database incident and from the Chubby and Physalia papers, what production systems actually do in the minutes when the instructions stop coming, and the one situation in which continuing to serve is the worst available move.

21 graded sources 8 production systems 4 published incidents Evidence through October 2026 Read: 27 min
01

The territory

The problem, stated without naming a technology: a large fleet does the work; a small strongly-consistent system tells it what the work is. What happens in the gap when the small system stops answering?

33 min
A Postgres primary kept committing after losing its leader lock, because nothing fenced it
50%
Envoy’s default panic threshold: the point at which it stops believing health checks
55%
Kubernetes’ unhealthy-zone threshold: above it, pod eviction slows 10× or stops
15%
Eureka’s self-preservation trigger: lose more leases than this at once and eviction halts

The split has many names (control plane and data plane, management server and proxy, distributed configuration store and cluster, registry and client) but one shape. The serving path handles every request. The instruction path tells the serving path who its members are, where traffic goes, which node holds an exclusive role. The two paths run at volumes that differ by orders of magnitude, and the instruction path almost always ends in a consensus system that, in etcd’s own words, serves “systems that will never tolerate split-brain operation and are willing to sacrifice availability to achieve this end”. That sentence, from etcd’s own design rationale, is the heart of the problem: the coordinator chooses consistency over availability on purpose, so every system downstream of it inherits an availability gap it must fill itself.

How systems fill that gap is unusually well documented in their repositories, because it is design-level behaviour that has to be written down for operators. Kubernetes wrote it as a founding principle: “components should continue to do what they were last told in the absence of new instructions.” Envoy wrote it into the xDS protocol: the last known configuration persists until the management server returns. Netflix wrote it into Eureka twice, once on the client (a registry cache that works “even when all of the eureka servers go down”) and once on the server (self-preservation mode). Patroni wrote the opposite: a primary that cannot renew its lock demotes itself immediately, because for a single-writer system continuing on stale authority is how you get two writers. The difference between those two answers, and the condition that selects between them, is the most useful thing this guide has to offer.

Scope. This guide covers the runtime dependency between a data plane and its coordinator: what happens from the moment the coordinator stops answering until it comes back. It deliberately excludes the change plane (deploys and config rollouts that propagate bad instructions rather than no instructions), client-side lock correctness and fencing tokens (a prior guide in this collection covers them), and consensus protocol internals. One honest limit up front: this session’s network reached code hosts only, so the AWS Builders’ Library essays on static stability, the Roblox 2021 Consul outage and the Cloudflare 2020 etcd outage, the three most cited texts in this territory, are absent. The mechanisms they describe appear here through artefacts that were reachable: code, design records, committed postmortems and GitLab’s public incident tracker.

Figure 1 · Two paths, two volumes, one dependency

every request
high volume

instructions: membership,
routes, leadership
low volume

heartbeats, lease renewals

Clients

Data plane
(proxies, kubelets,
replicas, agents)

Backends

Coordinator
(consensus store,
registry, xDS server)

every request
high volume

instructions: membership,
routes, leadership
low volume

heartbeats, lease renewals

Clients

Data plane
(proxies, kubelets,
replicas, agents)

Backends

Coordinator
(consensus store,
registry, xDS server)

Requests never touch the coordinator; instructions do. Losing the coordinator therefore starves the instruction path, not the serving path, and everything in this guide is about how long the serving path can run on stale instructions. Reconstructed from Kubernetes design principles and Envoy’s xDS protocol.
Diagram source
02

How it is actually built

Across Kubernetes, Envoy, Eureka, Patroni and the two Google and AWS papers, the same four mechanisms recur. Together they are the reference architecture for surviving the coordinator’s absence.

Figure 2 · The four mechanisms between coordinator and data plane

Data plane instance

Control plane

watch / push updates

health verdicts

lease renewals,
every ttl/3

4 · coordination only at
failure time, sized for bursts

Consensus store
CP by design

Controllers / management server

1 · Last-known-good cache
serve stale on silence

2 · Disbelief gate
ignore mass-death signals

3 · Lease + fence
exclusive roles expire and halt

Serving loop

Data plane instance

Control plane

watch / push updates

health verdicts

lease renewals,
every ttl/3

4 · coordination only at
failure time, sized for bursts

Consensus store
CP by design

Controllers / management server

1 · Last-known-good cache
serve stale on silence

2 · Disbelief gate
ignore mass-death signals

3 · Lease + fence
exclusive roles expire and halt

Serving loop

Every box is attributable: the last-known-good cache is Envoy’s xDS default and Eureka’s client cache; the disbelief gate is Envoy’s panic threshold, Kubernetes’ zone states and Eureka’s self-preservation; the lease and fence protect exclusive roles (Patroni, Chubby); failure-time-only coordination is Physalia’s design goal for EBS.
Diagram source

Mechanism one: the last-known-good cache. The data plane holds a complete local copy of its instructions and treats coordinator silence as “no change”. Envoy’s xDS protocol states it flatly: “in the event that the management server becomes unreachable, the last known configuration received by Envoy will persist until the connection is reestablished.” Eureka’s wiki makes the same promise for clients, which “can operate reasonably well, even when all of the eureka servers go down”, and names the price in the same breath: clients may be handed instances that no longer exist, so “the best protection in these scenarios is to timeout quickly and try other servers.” The cache does not remove the failure; it converts an outage into staleness, and staleness into a client-side retry problem. The transfer for your own design: a last-known-good cache is only as good as the data plane’s tolerance for acting on lies, and that tolerance must be engineered (fast timeouts, retries against alternatives), not assumed.

Mechanism two: the disbelief gate. Three systems built independently by three organisations each contain a numeric threshold above which the system stops believing its own health signal. Envoy: if more than 50% of a cluster’s hosts look unhealthy, disregard health status entirely. Kubernetes: if 55% of a zone’s nodes are NotReady (minimum three), evict at one-tenth speed, and in clusters of 50 nodes or fewer stop evicting altogether; and if every node looks dead, the controller logs “Controller detected that all Nodes are not-Ready. Entering master disruption mode”, marks every node reachable and swaps all eviction limiters to zero. Eureka: if more than 15% of the registry is pending eviction at once, stop evicting anything. None of the three documents cites the others. The sources give the shared reasoning: a signal that says one node died is probably reporting a dead node; a signal that says most nodes died at once is probably reporting a broken signal path, and acting on it would convert a connectivity problem into a workload massacre. I am calling this the disbelief threshold; the sources call it panic threshold, partial disruption and self-preservation, which is part of why nobody notices it is one idea.

Mechanism three: the lease, and the fence behind it. Exclusive roles cannot ride out coordinator loss on a cache, because the instruction “you are the only writer” becomes false the moment someone else can be elected. So leadership is leased with an explicit staleness budget: Patroni’s lock TTL defaults to 30 seconds, with the documented invariant loop_wait + 2 * retry_timeout <= ttl bounding how much checking fits inside one lease. The coordinator holds up its side: etcd’s failure-mode documentation promises that a new leader “extends timeouts automatically for all leases”, so an election pause does not cascade into mass lease expiry downstream. Chubby’s paper, twenty years old, states both the intent, “it is good for coarse-grained locks to survive lock server failures”, and the client contract, “clients must be prepared to lose locks during network partitions.” What the lease cannot do is make the deposed holder actually stop; that needs a fence, and section 4 shows what 33 minutes without one looks like.

Mechanism four: coordinate only at failure time, and size for the correlated case. The Physalia paper describes EBS as a “sometimes-coordinating” system: volumes replicate without consulting the configuration master during normal operation and reach for it only when a replica fails. The paper is equally clear about the trap in that shape: “recovery load is not constant, and highest during bad network conditions”, so the coordinator sees its heaviest traffic at exactly the moment infrastructure is least able to carry it. Physalia’s stated design target is worth stealing whole: optimise not the coordinator’s availability in general but the conditional probability that it is up given that a client needs it, which it writes as P(Av | Ai) and pursues by placing consensus cells physically near their clients.

Where each system keeps the cache

Envoy holds config in process memory; Eureka clients hold the full registry; kubelets hold pod specs and keep containers running regardless of API server state, which is the design principle quoted above doing its job.

Sources: xDS protocol, Eureka wiki, K8s principles

Where the thresholds live

All three disbelief gates ship as configurable constants: Envoy’s healthy_panic_threshold (50%), kube-controller-manager’s --unhealthy-zone-threshold (0.55), Eureka’s renewalPercentThreshold (0.85, i.e. act below 85% of expected renewals).

Sources: Envoy docs, KCM options, Eureka config

Where the divergence is

Read-and-route planes default to serving stale. Single-writer planes default to stopping: Patroni demotes to read-only the moment the lock update fails. The reference architecture is the same; the default flips with the writes.

Source: Patroni DCS failsafe docs

03

The decisions that matter

Four forks, each with the condition that flips it. The first one is the spine of the whole topic.

Decision 1: when the coordinator goes silent, does the data plane keep acting on its last instructions?

Chosen
  • Keep serving on last-known-good: Envoy (xDS default), Eureka clients, kubelet, by explicit design principle
  • Because instruction loss is far more common than instruction change, and an idle fleet helps nobody
Rejected
  • Fail closed: drop config, stop serving until the coordinator returns
  • Rejected for routing data because it converts every coordinator blip into a customer-visible outage
Flips when
  • The instruction confers exclusive authority (a write leadership): Patroni demotes immediately because stale authority means two writers
  • The instruction is dangerous when stale: Envoy’s docs name fault-injection rules, and provide a TTL so they expire on contact loss instead of persisting

Decision 2: when the health signal reports mass death, is it believed?

Chosen
  • A disbelief threshold: 50% (Envoy), 55% (Kubernetes), 15% (Eureka); above it, freeze evictions or ignore health status
  • Because a correlated death report usually means the reporting path broke, not the fleet
Rejected
  • Trusting each signal individually at any scale of failure
  • Eureka’s wiki names the rejected outcome: catastrophic network events wiping the registry and propagating emptiness to every client
Flips when
  • The upstream genuinely fails all-or-nothing: Envoy documents fail_traffic_on_panic for services where sending traffic to a mostly-dead cluster just deepens the hole
  • Then failing closed during panic is the better trade, and Envoy states the condition in its own docs

Decision 3: a write leader loses the coordinator. Demote on suspicion, or verify with peers?

Chosen
  • Patroni default: demote to read-only immediately, since lock failure is indistinguishable from a partition from one node’s view
  • Failsafe mode (opt-in, Patroni 3.0): stay primary only if every known member acknowledges you over REST, re-checked every loop
Rejected
  • Keeping the primary on a quorum of members. Patroni’s docs reject it by name: the DCS’s quorum and the database cluster’s quorum are different views, and they can disagree about which side of a partition won
Flips when
  • Availability of writes matters more than the risk window and membership is small and static: unanimity is then cheap, and failsafe mode earns its place
  • The DCS shares failure domains with the database network (a Kubernetes API backed by the same etcd), which raises how often the default needlessly demotes

Decision 4: does the coordinator sit in the request path, or only in the failure path?

Chosen
  • Failure path only. EBS volumes replicate without the control plane and consult it when a replica dies (Physalia); Eureka pushes discovery into client caches instead of a proxy hop
Rejected
  • Consulting shared control-plane state at high rate. The 2011 EBS outage is the recorded argument: re-mirroring negotiations hammered a control plane that shared a database with API traffic, and the brownout spread region-wide
Flips when
  • Instructions must be revocable in seconds (authorisation, kill switches): then the consult is the product, and the coordinator’s availability becomes the product’s, to be bought with replication and cells rather than caching

Figure 3 · What should this data plane do when the coordinator is gone?

yes

no

yes

no

Coordinator
unreachable

Does the instruction grant
an exclusive role?

Demote / stop within the lease TTL
(Patroni default: immediately)

Back it with a fence:
watchdog or self-kill

Is the instruction dangerous
when stale? (faults, grants)

Attach a TTL so it expires
on contact loss (xDS TTL)

Serve last-known-good and
alert on staleness age

yes

no

yes

no

Coordinator
unreachable

Does the instruction grant
an exclusive role?

Demote / stop within the lease TTL
(Patroni default: immediately)

Back it with a fence:
watchdog or self-kill

Is the instruction dangerous
when stale? (faults, grants)

Attach a TTL so it expires
on contact loss (xDS TTL)

Serve last-known-good and
alert on staleness age

Terminal nodes are actions. The left branch is Patroni’s world, the middle is Envoy’s TTL exception, the right is the xDS and Eureka default; the dashed box is the fence that section 4 shows failing.
Diagram source
DecisionChosenRejectedBecauseEvidence
Config on coordinator lossPersist last-known-goodDrop and fail closedSilence is more common than changexDS protocol
Stale fault-injection rulesExpire via TTLPersist like routesA stuck fault is an outage you injectedxDS protocol, TTL section
Mass-unhealth signalDisbelief thresholdBelieve every signalCorrelated death usually means broken signal pathnode_lifecycle_controller.go
Panic behaviourFail open across all hostsFail closedSome successes beat none, unless failure is all-or-nothingEnvoy panic docs
Primary on DCS lossDemote now; failsafe needs unanimityQuorum of membersDCS quorum and member quorum are different viewsPatroni failsafe docs
Coordinator placementFailure path onlyIn the request path2011 EBS brownout via shared databasePhysalia paper
04

What broke in production

Three failure classes cover the published record: the half-dead holder, the coordinator in the request path, and the coordinator whose own state is wrong.

Figure 4 · INC-14518: the half-dead primary, step by step

Replica (node 03)Consul serversPostgres (samenode)Consul + PatroniagentsGCP boot diskReplica (node 03)Consul serversPostgres (samenode)Consul + PatroniagentsGCP boot diskHA loop andmembership frozenlogs "Could notactivate Linuxwatchdog device"two writable timelines divergeI/O stall, effectively unboundedtimeoutlease renewals stopleader lock expires(TTL)lock free, node 03 promoteskeeps committing ~33min, unfenced
Replica (node 03)Consul serversPostgres (samenode)Consul + PatroniagentsGCP boot diskReplica (node 03)Consul serversPostgres (samenode)Consul + PatroniagentsGCP boot diskHA loop andmembership frozenlogs "Could notactivate Linuxwatchdog device"two writable timelines divergeI/O stall, effectively unboundedtimeoutlease renewals stopleader lock expires(TTL)lock free, node 03 promoteskeeps committing ~33min, unfenced
The control agents and the database shared a node but not a fate: the disk stall froze Consul and Patroni while Postgres kept committing. Sequence reconstructed from GitLab’s public corrective actions and runbook.
Diagram source
Postmortem

The half-dead primary: GitLab INC-14518

AssumptionA node that loses its leader lock stops writing; demotion and fencing are the same thing.
What happenedA GCP persistent-disk I/O stall “froze Consul, Patroni’s HA loop, and other processes relying on the boot disk” while Postgres, on a different disk, kept running. The lock expired; nothing fenced the old primary; it “kept committing for about 33 minutes.”
Blast radiusTwo writable timelines on the main production database of GitLab.com; recovery required in-place assessment rather than failover, since the stuck node held the most complete data.
FixLoad a software watchdog at boot, set Patroni’s watchdog mode to required (a node “will not become a leader unless watchdog can be successfully enabled”), add leaderless-cluster alerts fed by Consul’s catalog, and alert when a cluster is left paused.
Design ruleFail-static is the correct default for everything except exclusivity. Wherever the data plane can keep acting after losing its role, demotion must be enforced by something that does not share the agent’s failure mode: a kernel watchdog, not another userspace loop on the same stalled disk.
Paper

The coordinator in the request path: EBS, April 2011

AssumptionThe control plane can share a database with API traffic because recovery traffic is small.
What happenedA network change left 13% of one AZ’s EBS volumes unable to re-mirror; each required a negotiation with the control plane, retries multiplied the calls, and “the load caused a brown out of the EBS control plane and again affected EBS APIs across the Region.”
Blast radiusRegional API impact from a single-AZ data-plane fault: the dependency ran backwards, data plane load taking down the control plane for everyone.
FixPhysalia: consensus split into millions of per-volume cells, placed near their clients, with blast radius as the headline design goal and the control plane out of the steady-state path.
Design ruleCoordination demand is a spike correlated with your worst day. Size the coordinator for the correlated failure case or keep it out of any path that scales with failures; the paper calls this the inherent risk of sometimes-coordinating systems.
Postmortem

The coordinator’s state is wrong: etcd v3.5 inconsistency

AssumptionThe consensus store underneath everything is the one component that does not lie.
What happenedA v3.5.0 refactor stopped saving the consistent index atomically with its transaction, so a crash between the two left members with diverging state: “independent crash could lead to committed transactions are not reflected on all the members.”
Blast radiusA year in the wild (v3.5.0 to v3.5.2, mid-2021 to April 2022) before the public statement; no confirmed production corruption reported, but the maintainers note the corruption check itself “might fail causing the check to pass,” and a single-member cluster’s divergence is “totally undetectable.”
Fixv3.5.3/v3.5.4 made the write atomic again; corruption detection was hardened and a robustness-testing programme followed.
Design ruleBudget for the coordinator being wrong, not just down: enable its integrity checks, keep restorable backups of its state, and rehearse the restore, because detection cannot be assumed to fire.
Postmortem

The coordinator’s membership is wrong: etcd zombie members

AssumptionRemoving a member removes it; membership is the one record consensus keeps perfectly.
What happenedOld clusters carried silent divergence between etcd’s two membership stores; upgrading to v3.6 flipped the source of truth, and members “removed from the database cluster some time ago” reappeared “and joining database consensus. The etcd cluster is then inoperable until these zombie members are removed.”
Blast radiusAffected clusters are inoperable mid-upgrade; the maintainers could offer no safe workaround other than reaching v3.5.26 first, and the trigger set (old snapshot restores, force-new-cluster recoveries) is precisely the clusters with the most operational history.
Fixv3.5.26 auto-syncs the two stores before the upgrade path flips the authority.
Design ruleEvery past disaster recovery on the coordinator is a liability on its next upgrade. Record which clusters were ever force-recovered or snapshot-restored, and treat coordinator upgrades on those as migrations needing verification, not routine rollouts.
The pattern across all four

No incident here is “the coordinator went down and everything stopped.” The data planes largely did their job. The damage came from the edges of the contract: a fence that was configuration rather than hardware, a coordinator borrowed into the hot path, and twice the coordinator’s own state going quietly wrong. The dependency you drew on the architecture diagram is rarely the one that gets you; its silent assumptions are.

05

Numbers you can plan against

Every figure is a shipped default, a measured value or a documented limit, with its date. Defaults are what most installations actually run.

MetricValueAtContextAs ofSource
Heartbeat interval100 msetcdDefault; set to ~RTT between membersv3.5 docs, 2026tuning.md
Election timeout1,000 msetcdDefault; documented ceiling 50 s for global clustersv3.5 docs, 2026tuning.md
Leader election pause4–6 s, up to 30 sGoogle ChubbyMeasured in production; writes queue during election2006Chubby paper
Leader lock TTL30 sPatroniDefault staleness budget before failover starts; invariant loop_wait + 2×retry_timeout ≤ ttldocs, 2026dynamic_configuration.rst
Watchdog fire point85 sGitLab gprdPatroni TTL minus safety_margin after last successful renewal, their settings2026-10runbook
Unfenced divergence window~33 minGitLab INC-14518Measured: primary kept committing after losing its lock, no watchdog2026-09-24corrective action
Panic threshold50%EnvoyDefault; below it, health status is disregardeddocs, 2026panic_threshold.rst
Unhealthy-zone threshold55% (min 3 nodes)KubernetesDefault; eviction drops from 0.1 to 0.01 nodes/s, and to 0 at ≤50 nodesmaster, 2026KCM options
Self-preservation trigger>15% evictions pending / renewals <85% in 15 minNetflix EurekaDefault; server stops expiring instances entirelywiki, 2026Eureka wiki
Registry warm-up hold5 minNetflix EurekaA server that cannot sync peers waits before serving, to avoid handing out a partial registrywiki, 2026Eureka wiki
Volumes impaired by control-plane event13% of one AZAWS EBSMeasured, April 2011; the incident behind Physalia2011/2020Physalia paper
Control-plane cost per remote cluster1% CPU core, 180 MBIstio ambientVendor load test, 300 services × 4,000 endpoints per clusterdocs, 2026Istio docs
Read these carefully

The etcd, Patroni, Envoy, Kubernetes and Eureka rows are shipped defaults, not measurements of your environment; they tell you what the ecosystem considers a sane staleness budget (tens of seconds for leases, minutes for registries), and they move with version upgrades. The Chubby and EBS figures are measured but dated 2006 and 2011; treat them as orders of magnitude. The Istio figure is a vendor’s own load test. The one number to internalise is the ratio: detection budgets are seconds, but the GitLab window shows an unfenced failure running 60× past its 30-second budget.

Figure 5 · The modes a data-plane instance moves through

coordinator unreachable

reconnect, resync

TTL hit (dangerous-when-stale config)

exclusive role, lease expires

watchdog or self-kill confirms the stop

exclusive role, no fence

Normal

Stale

Expired

Demoted

Fenced

Split authority (two writers)

GitLab INC-14518 spent
33 minutes here

coordinator unreachable

reconnect, resync

TTL hit (dangerous-when-stale config)

exclusive role, lease expires

watchdog or self-kill confirms the stop

exclusive role, no fence

Normal

Stale

Expired

Demoted

Fenced

Split authority (two writers)

GitLab INC-14518 spent
33 minutes here

Operators see these as dashboard states; the dangerous transition is the dashed one, where an exclusive role rides out coordinator loss without a fence. Mapped from the mechanisms in sections 2 and 4.
Diagram source
06

The evidence wall

Every source behind this page, graded. This session’s network reached code hosts only, so the wall is repository artefacts, GitLab’s public tracker, and two paper mirrors; there are no talks, and the engineering-blog layer is thin by constraint rather than choice.

Postmortem GitLab2026-09

INC-14518 corrective action: Patroni fencing and self-shutdown

The public record of the half-dead primary: a GCP disk stall froze Consul and Patroni, the node lost its lock, and Postgres kept committing for about 33 minutes because the watchdog was never configured and /dev/watchdog did not exist.

Carry forwardDemotion is a decision; fencing is a mechanism. Only the mechanism counts when the agents are frozen.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/23046
Postmortem GitLab2026-10

INC-14518 corrective actions: discovery path and paused clusters

Two more public follow-ups: the Patroni/Consul/pgbouncer discovery path was undocumented, which slowed diagnosis; and a paused Patroni cluster (no failover, no watchdog) raised no alert at all.

Carry forwardA disabled control loop is a standing incident with no alarm. Alert on the loop’s state, not just its outputs.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/23070
Postmortem etcd maintainers2022-04

v3.5 data inconsistency postmortem

Committed to the repository itself: a refactor broke the atomicity of the consistent index, members could diverge after crashes, and the corruption check depends on a call that can fail in a way that passes the check.

Carry forwardThe coordinator is software. Run its integrity checks and rehearse restoring it; do not model it as a constant.
github.com/etcd-io/etcd/.../v3.5-data-inconsistency.md
Postmortem etcd maintainers2025-12

Avoiding zombie cluster members when upgrading to v3.6

Years-old divergence between etcd’s two membership stores resurfaced as removed members rejoining consensus on upgrade, leaving clusters inoperable; the triggers are old disaster-recovery actions.

Carry forwardEvery force-recovery of the coordinator leaves a scar; inventory them before the next coordinator upgrade.
etcd.io/blog/2025/zombie_members_upgrade
Decision record Kubernetespre-2018

Design principles: continue on last instructions

The founding architecture principles include the whole pattern in one line: components should continue to do what they were last told in the absence of new instructions, with partition and component outage named as the motivating cases.

Carry forwardMake staleness tolerance a stated principle, so every new controller inherits it instead of re-deciding it.
github.com/kubernetes/design-proposals-archive/.../principles.md
Decision record Envoychecked 2026

xDS protocol: last-known-good, and the TTL exception

The protocol document states the fail-static default and then the exception that proves the rule: fault-injection config left behind by a dead management server is dangerous, so resources can carry a TTL that removes them on contact loss.

Carry forwardClassify every config kind as safe-stale or dangerous-stale, and attach expiry only to the second.
github.com/envoyproxy/envoy/.../xds_protocol.rst
Decision record Patronichecked 2026

DCS failsafe mode: the rejected quorum, in writing

The design document for what a primary may do when the DCS is gone: demote by default, survive only on unanimous member acknowledgement, and an FAQ that rejects quorum because the DCS and the database cluster can hold different views of the same partition.

Carry forwardWhen two systems each have a quorum, their intersection is not one; unanimity is the only cheap safe check.
github.com/patroni/patroni/blob/master/docs/dcs_failsafe_mode.rst
Decision record Kubernetes2015-12

Control plane resilience design (Hoole, Danese, Santa-Barbara)

Separates self-healing from high availability and notes that replicas without replacement “fairly obviously” drift to unavailability over time; the control plane needs both properties, independently engineered.

Carry forwardAn HA coordinator that nothing repairs is a slow outage; budget the repair loop, not just the replica count.
github.com/kubernetes/design-proposals-archive/.../control-plane-resilience.md
Source code Kubernetesmaster, 2026

node_lifecycle_controller.go: zone states and master disruption mode

The disbelief gate in executable form: ComputeZoneState flips a zone to partialDisruption at 55% NotReady, and when every zone is fully disrupted the controller marks all nodes reachable and swaps every eviction limiter to zero.

Carry forwardRate-limit destructive reactions by the fraction of the fleet affected, with a hard stop at “everyone looks dead.”
github.com/kubernetes/kubernetes/.../node_lifecycle_controller.go
Source code Kubernetesmaster, 2026

kube-controller-manager node lifecycle flags

The shipped numbers with their own documentation: eviction at 0.1 nodes/s when healthy, 0.01 when a zone is unhealthy, overridden to zero below 50 nodes, threshold 0.55 with a minimum of three NotReady nodes.

Carry forwardSmall clusters get no secondary eviction at all; upstream decided that for you, and it is probably right.
github.com/kubernetes/kubernetes/.../options/nodelifecyclecontroller.go
Source code Netflixmaster, 2026

Eureka: PeerAwareInstanceRegistryImpl and server config

Lease expiry runs only while renewals exceed the threshold (renewalPercentThreshold 0.85); the registry would rather serve dead entries than evict live ones on a broken signal.

Carry forwardExpiry driven by heartbeats must be gated on the heartbeat channel itself looking healthy.
github.com/Netflix/eureka/.../PeerAwareInstanceRegistryImpl.java
Source code GitLab2026-10

Runbook: primary unresponsive, recover without losing data

Operational doctrine written after INC-14518: do not fail over first, because the half-dead primary is the most complete copy of the data; with their settings the watchdog fires 85 seconds after the last successful lock renewal; assessment is timeboxed to 10 minutes.

Carry forwardWrite the “half-dead” branch into the runbook before the incident; it is the branch responders get wrong under pressure.
gitlab.com/gitlab-com/runbooks/.../primary-unresponsive-assess-before-failover.md
Source code GitLab2026-10, open

Runbooks MR 11682: leaderless-cluster alerts that guard their own inputs

The proposed alerts detect a leaderless Patroni cluster from Consul’s catalog, and a third, non-paging alert fires when the inputs of the first two go absent, because an alert fed by the control plane disappears with it.

Carry forwardAny alert whose data flows through the coordinator needs an absent() guard, or the outage silences its own alarm.
gitlab.com/gitlab-com/runbooks/-/merge_requests/11682
Paper Google (Burrows)2006-11

The Chubby lock service (OSDI ’06; PDF mirror)

The origin text for coarse-grained coordination that survives its own failures: five replicas per cell, elections measured at 4 to 6 seconds with a 30-second tail, and the client contract that locks can vanish in partitions without creating new recovery paths.

Carry forwardDesign clients so losing coordination is a slow path they already have, not a new code path discovered in production.
raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf
Paper AWS (Brooker, Chen, Ping)2020-02

Millions of Tiny Databases (NSDI ’20; PDF mirror)

Physalia: consensus for EBS re-sharded into per-volume cells after the 2011 outage, designed to maximise the probability the coordinator is up conditional on a client needing it, with the warning that recovery load peaks during bad network conditions.

Carry forwardMeasure your coordinator’s availability conditioned on failure events, not averaged over calendar time.
raw.githubusercontent.com/arpit20adlakha/.../Millions of Tiny Databases.pdf
Eng blog Netflixchecked 2026

Eureka wiki: self-preservation and peer-to-peer behaviour

Netflix-authored operational doctrine: stop evicting when more than 15% of the registry looks dead at once, hold a cold server for five minutes rather than serve a partial registry, and tell clients plainly that they must tolerate stale entries.

Carry forwardA registry that can be emptied by a network event will be; prefer stale-and-full to fresh-and-empty.
github.com/Netflix/eureka/wiki/Server-Self-Preservation-Mode
Eng blog Netflixchecked 2026

Eureka at a glance: the client cache as the availability story

States the architecture bet directly: discovery data cached in every client means applications resilient to the outage of the registry itself, contrasted explicitly with proxy-based load balancing where the balancer’s outage is yours.

Carry forwardMoving a dependency out of the request path and into a cached instruction is the cheapest nine you will ever buy.
github.com/Netflix/eureka/wiki/Eureka-at-a-glance
Vendor docs Envoychecked 2026

Panic threshold, with the fail-open/fail-closed switch

Official docs for the 50% default, the two panic behaviours, and the stated condition for choosing fail-closed: upstreams that fail all-or-nothing, where partial traffic only deepens the failure.

Carry forwardDecide per upstream whether some success beats fast failure; the default assumes it does.
github.com/envoyproxy/envoy/.../panic_threshold.rst
Vendor docs etcdv3.5, checked 2026

Failure modes, tuning, and why etcd

The coordinator’s own contract: writes pause during elections but committed writes survive; leases are auto-extended across elections; majority loss stops writes entirely; and the explicit statement that availability is sacrificed for split-brain safety.

Carry forwardRead your coordinator’s failure-modes page as an SLA; everything it declines to promise is your staleness budget.
github.com/etcd-io/website/.../op-guide/failures.md
Vendor docs Istiochecked 2026

Deployment models and control-plane cost

Control-plane scope treated as outage scope (more control planes, smaller blast radius), plus a vendor load test pricing a remote cluster at roughly 1% of a CPU core and 180 MB of control-plane memory at 300 services and 4,000 endpoints.

Carry forwardControl planes are cheap enough to multiply; blast radius, not cost, should set their count.
github.com/istio/istio.io/.../deployment-models/index.md
Vendor docs Patronichecked 2026

Dynamic configuration: the lease arithmetic

TTL 30 seconds by default, minimum 20, with the invariant loop_wait + 2×retry_timeout ≤ ttl that bounds how many verification attempts fit inside one lease; member slots retained 30 minutes after a member vanishes.

Carry forwardYour failover time is a sum you can compute from config; compute it before an incident does.
github.com/patroni/patroni/blob/master/docs/dynamic_configuration.rst
07

Build a miniature, then productionise it

Six rungs from an evening toy to a tested staleness budget on your real system. The crossing from toy to real is rung four.

Kill the management server

Run Envoy against a minimal xDS server (a static file server speaking SotW is enough), send steady traffic, then kill the server.

Done when: traffic flows unchanged for ten minutes with the management server dead, and you can say where the config lives in the meantime.  Teaches: fail-static as the shipped default, not a feature you add.

Trip the panic threshold

With the same setup, fail health checks on 60% of upstream hosts. Watch Envoy disregard health status at the 50% default; then set fail_traffic_on_panic and repeat.

Done when: you can state, for one real service you run, which panic mode it wants and why.  Teaches: the disbelief threshold and the fail-open/fail-closed trade.

Demote a primary by stopping its coordinator

Three-node Patroni on etcd. Stop etcd entirely and time the demotion against the 30-second TTL. Enable failsafe_mode and repeat, watching the primary survive on unanimous member acknowledgement.

Done when: you have both timelines measured and can explain why quorum of members would be unsafe.  Teaches: the staleness budget as lease arithmetic.

Make the leader half-dead

GitLab’s own validation step: kill -STOP the Patroni process on the leader under write load, leaving Postgres running. Measure how long writes continue after the lock expires. Then load softdog, set watchdog mode to required, and repeat.

Done when: the watchdog resets the frozen leader before the lock expires, and you have counted the writes lost in each configuration.  Teaches: why fail-static plus exclusivity demands a fence outside the process.

Build a registry that survives a partition

Write a toy registry that expires entries after three missed heartbeats. Partition 30% of clients away and watch it evict them all. Add Eureka’s rule: if expected renewals drop below 85%, stop expiring.

Done when: the partition no longer empties the registry, and stale entries are visibly flagged with their age.  Teaches: acting on absence, and why the threshold must gate the eviction loop itself.

Run the game day on your real system

Inventory every runtime read your serving path makes against a coordinator (service discovery, feature flags, leases, config watches). Block the coordinator in staging for 30 minutes and record what degrades, when, and what the dashboards claimed.

Done when: a written staleness budget per dependency, signed off by the owning team, with the first alert firing on staleness age rather than coordinator death.  Teaches: the dependency list is always longer than the diagram, and recovery load is part of the budget.

08

Keep hunting

The queries and commands that found this material, in the forms that worked. The repository layer is unusually rich for this topic because staleness behaviour must be documented for operators.

Incident trackers and postmortems in repos

  • curl "https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=consul&state=all"
  • curl "https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=patroni+failover&labels=Incident::Resolved"
  • etcd postmortem path:Documentation/postmortems
  • "corrective action" watchdog fencing site:gitlab.com

Repository archaeology

  • git clone --depth 1 --filter=blob:none --sparse https://github.com/kubernetes/kubernetes
  • grep -rn "unhealthyZoneThreshold\|SwapLimiter(0)" pkg/controller/nodelifecycle/
  • grep -rn "last known configuration" docs/root/api-docs/xds_protocol.rst
  • grep -rln "failsafe\|watchdog" patroni/docs/
  • grep -rn "renewalPercentThreshold" eureka-core/

The design-record layer

  • "continue to do what they were last told" kubernetes principles
  • repo:kubernetes/design-proposals-archive resilience OR disruption
  • "eventual consistency considerations" xds
  • "failsafe_mode" patroni "all known members"

Vocabulary that unlocks the field

  • "static stability" control plane data plane
  • "panic threshold" OR "self preservation" OR "unhealthy-zone-threshold"
  • "sometimes-coordinating" OR "blast radius" consensus cells
  • "half-dead" OR "stuck but not down" primary postmortem
09

References

  1. Kubernetes, Design principles kubernetes/design-proposals-archive, pre-2018 archive. Checked 2026-10-06.
  2. Hoole, Danese, Santa-Barbara, Kubernetes and Cluster Federation Control Plane Resilience kubernetes/design-proposals-archive, 2015-12-14. Checked 2026-10-06.
  3. Kubernetes, node_lifecycle_controller.go kubernetes/kubernetes master. Checked 2026-10-06.
  4. Kubernetes, kube-controller-manager node lifecycle options kubernetes/kubernetes master. Checked 2026-10-06.
  5. Envoy, xDS REST and gRPC protocol envoyproxy/envoy main. Checked 2026-10-06.
  6. Envoy, Panic threshold envoyproxy/envoy main. Checked 2026-10-06.
  7. Netflix, Eureka wiki: Server Self Preservation Mode Netflix/eureka wiki. Checked 2026-10-06.
  8. Netflix, Eureka wiki: Understanding Eureka Peer to Peer Communication Netflix/eureka wiki. Checked 2026-10-06.
  9. Netflix, Eureka wiki: Eureka at a glance Netflix/eureka wiki. Checked 2026-10-06.
  10. Netflix, PeerAwareInstanceRegistryImpl.java Netflix/eureka master. Checked 2026-10-06.
  11. Netflix, DefaultEurekaServerConfig.java Netflix/eureka master. Checked 2026-10-06.
  12. Patroni, DCS Failsafe Mode patroni/patroni master. Checked 2026-10-06.
  13. Patroni, Dynamic configuration patroni/patroni master. Checked 2026-10-06.
  14. GitLab, Corrective action: Bugfix Patroni fencing and self-shutdown (INC-14518) gl-infra/production tracker, 2026-09-29. Checked 2026-10-06.
  15. GitLab, Corrective action: review failover candidate policy and document discovery path (INC-14518) gl-infra/production tracker, 2026-10-01. Checked 2026-10-06.
  16. GitLab, Corrective action: alert when a Patroni cluster is left paused (INC-14518) gl-infra/production tracker, 2026-10-01. Checked 2026-10-06.
  17. GitLab, Runbook: Primary unresponsive or no Leader gitlab-com/runbooks master. Checked 2026-10-06.
  18. GitLab, Runbooks MR 11682: Patroni leaderless-cluster and replica-streaming alerts gitlab-com/runbooks, open at check time. Checked 2026-10-06.
  19. etcd maintainers, v3.5 data inconsistency postmortem etcd-io/etcd, 2022-04-20. Checked 2026-10-06.
  20. Wang, Berkus, Avoiding Zombie Cluster Members When Upgrading to etcd v3.6 etcd.io blog, 2025-12-17; markdown source read from etcd-io/website. Checked 2026-10-06.
  21. etcd, Failure modes etcd-io/website (docs v3.5). Checked 2026-10-06.
  22. etcd, Why etcd etcd-io/website (docs v3.5). Checked 2026-10-06.
  23. etcd, Tuning etcd-io/website (docs v3.5). Checked 2026-10-06.
  24. Burrows, The Chubby lock service for loosely-coupled distributed systems (OSDI ’06) PDF mirror on GitHub (canonical: research.google archive). Checked 2026-10-06.
  25. Brooker, Chen, Ping, Millions of Tiny Databases (NSDI ’20) PDF mirror on GitHub (canonical: USENIX NSDI ’20). Checked 2026-10-06.
  26. Istio, Deployment models istio/istio.io master. Checked 2026-10-06.
  27. Istio, Ambient multicluster performance istio/istio.io master. Checked 2026-10-06.