Evidence ledger
One row per claim in Serving without the control plane: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Topic: how data planes are designed to keep serving while the control plane that instructs them is down, and when continuing to serve is exactly the wrong move.
Network constraint for this session: outbound access reached code hosts only (github.com, raw.githubusercontent.com, gitlab.com and the GitLab API, pkg.go.dev). Every repository file cited below was read in full from a clone made in this session; GitLab issues, work items, merge requests and raw files were fetched over the GitLab API or raw endpoints. The two papers were read from PDF mirrors hosted in public GitHub repositories and fetched in this session. etcd.io blog URLs are cited canonically; their full markdown source was read from the etcd-io/website repository clone. Engineering blogs, talk recordings and postmortems hosted anywhere else (AWS, Roblox, Cloudflare, HashiCorp, USENIX) were unreachable and are therefore absent, not overlooked.
All links checked 2026-10-06. One row per claim. Quotes are copied, not paraphrased.
| # | Org | Title | Tier | Published | Checked | URL | Claim taken from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | Kubernetes | Design principles (design-proposals-archive) | adr | pre-2018 archive, undated | 2026-10-06 | https://github.com/kubernetes/design-proposals-archive/blob/main/architecture/principles.md | The whole pattern stated as a design principle | "Components should continue to do what they were last told in the absence of new instructions (e.g., due to network partition or component outage)." |
| 2 | Envoy | xDS REST and gRPC protocol | adr | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/envoyproxy/envoy/blob/main/docs/root/api-docs/xds_protocol.rst | The proxy's default on management-server loss is to keep the last configuration | "In the event that the management server becomes unreachable, the last known configuration received by Envoy will persist until the connection is reestablished." |
| 3 | Envoy | xDS REST and gRPC protocol | adr | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/envoyproxy/envoy/blob/main/docs/root/api-docs/xds_protocol.rst | The documented exception: configs that are dangerous when stale get a TTL | "For some services, this may not be desirable. For example, in the case of a fault injection service, a management server crash at the wrong time may leave Envoy in an undesirable state. The TTL setting allows Envoy to remove a set of resources after a specified period of time if contact with the management server is lost." |
| 4 | Envoy | Panic threshold (arch overview) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/envoyproxy/envoy/blob/main/docs/root/intro/arch_overview/upstream/load_balancing/panic_threshold.rst | At 50% unhealthy, Envoy stops believing health checks | "if the percentage of available hosts in the cluster becomes too low, Envoy will disregard health status and balance either amongst all hosts or no hosts. This is known as the panic threshold. The default panic threshold is 50%." |
| 5 | Envoy | Panic threshold (arch overview) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/envoyproxy/envoy/blob/main/docs/root/intro/arch_overview/upstream/load_balancing/panic_threshold.rst | Fail-open vs fail-closed is an explicit configuration, with the trade stated | "traffic will either be sent to all hosts, or will be sent to no hosts (and therefore will always fail) ... Choosing to fail traffic during panic scenarios can help avoid overwhelming potentially failing upstream services ... However, it eliminates the possibility of some requests succeeding even when many or all hosts in a cluster are unhealthy." |
| 6 | Kubernetes | node_lifecycle_controller.go | source | current master, retrieved 2026-10-06 | 2026-10-06 | https://github.com/kubernetes/kubernetes/blob/master/pkg/controller/nodelifecycle/node_lifecycle_controller.go | A zone is declared disrupted at the 55% threshold, minimum 3 nodes | "case notReadyNodes > 2 && float32(notReadyNodes)/float32(notReadyNodes+readyNodes) >= nc.unhealthyZoneThreshold: return notReadyNodes, statePartialDisruption" |
| 7 | Kubernetes | node_lifecycle_controller.go | source | current master, retrieved 2026-10-06 | 2026-10-06 | https://github.com/kubernetes/kubernetes/blob/master/pkg/controller/nodelifecycle/node_lifecycle_controller.go | When every node looks dead, the controller distrusts the signal entirely and stops evicting | "logger.Info(\"Controller detected that all Nodes are not-Ready. Entering master disruption mode\") ... // We stop all evictions. for k := range nc.zoneStates { nc.zoneNoExecuteTainter[k].SwapLimiter(0) }" — and it calls markNodeAsReachable on every node. |
| 8 | Kubernetes | kube-controller-manager node lifecycle options | source | current master, retrieved 2026-10-06 | 2026-10-06 | https://github.com/kubernetes/kubernetes/blob/master/cmd/kube-controller-manager/app/options/nodelifecyclecontroller.go | The shipped defaults: 0.55 threshold, 10x slower eviction when disrupted, zero below 50 nodes | "--node-eviction-rate" default 0.1; "--secondary-node-eviction-rate" default 0.01, "implicitly overridden to 0 if the cluster size is smaller than --large-cluster-size-threshold" (default 50); "--unhealthy-zone-threshold" default 0.55, "Fraction of Nodes in a zone which needs to be not Ready (minimum 3) for zone to be treated as unhealthy." |
| 9 | Netflix | Eureka wiki: Server Self Preservation Mode | blog | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/Netflix/eureka/wiki/Server-Self-Preservation-Mode | The registry stops evicting when >15% of instances look dead at once | "it is when > 15% of the current registry is in this later state, that self preservation will be enabled. When in self preservation mode, eureka servers will stop eviction of all instances" |
| 10 | Netflix | Eureka wiki: Understanding Eureka Peer to Peer Communication | blog | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/Netflix/eureka/wiki/Understanding-Eureka-Peer-to-Peer-Communication | The threshold in operation (85% of expected renewals over 15 minutes) and its stated price: stale instances, which the client must absorb | "If any time, the renewals falls below the percent configured for that value (below 85% within 15 mins), the server stops expiring instances to protect the current instance registry information. ... this may cause the clients to get the instances that do not exist anymore. The clients must make sure they are resilient to eureka server returning an instance that is non-existent or un-responsive." |
| 11 | Netflix | Eureka wiki: Eureka at a glance | blog | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/Netflix/eureka/wiki/Eureka-at-a-glance | The client cache is the availability mechanism: serving survives total registry loss | "Since Eureka clients have the registry cache information in them, they can operate reasonably well, even when all of the eureka servers go down." |
| 12 | Netflix | DefaultEurekaServerConfig.java | source | current master, retrieved 2026-10-06 | 2026-10-06 | https://github.com/Netflix/eureka/blob/master/eureka-core/src/main/java/com/netflix/eureka/DefaultEurekaServerConfig.java | The 85% renewal threshold is a shipped constant | "renewalPercentThreshold", 0.85 |
| 13 | Patroni | DCS Failsafe Mode (docs) | adr | Patroni 3.0 feature, doc retrieved 2026-10-06 | 2026-10-06 | https://github.com/patroni/patroni/blob/master/docs/dcs_failsafe_mode.rst | The default for a write leader on DCS loss is immediate demotion, because loss is indistinguishable from partition | "the node is allowed to run Postgres as the primary only if it can update the leader lock in DCS. In case the update of the leader lock fails, Postgres is immediately demoted and started as read-only. ... In general, it is impossible to distinguish between these two from a single node, and therefore Patroni assumes the worst case - network partitioning." |
| 14 | Patroni | DCS Failsafe Mode (docs) | adr | Patroni 3.0 feature, doc retrieved 2026-10-06 | 2026-10-06 | https://github.com/patroni/patroni/blob/master/docs/dcs_failsafe_mode.rst | The engineered middle: keep the primary only if every member acknowledges it, and the quorum alternative was rejected with a reason | "Postgres may continue to run as a primary if it can access all known members of the cluster via Patroni REST API ... Why MUST the current primary see ALL other members? Can't we rely on quorum here? ... The problem is that the view on the quorum might be different from the perspective of DCS and Patroni." |
| 15 | GitLab | Corrective action: Bugfix Patroni fencing and self-shutdown (INC-14518) | postmortem | 2026-09-29 | 2026-10-06 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23046 | A primary that lost its lock kept committing for 33 minutes because nothing fenced it | "On Sep 24 nothing fenced patroni-main-v17-01 after it lost the leader lock. The node kept committing for about 33 minutes ... patroni.yml has no watchdog: block ... softdog is never loaded, so /dev/watchdog doesn't exist." |
| 16 | GitLab | Corrective action: Bugfix Patroni fencing and self-shutdown (INC-14518) | postmortem | 2026-09-29 | 2026-10-06 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23046 | The trigger froze the control agents while the data plane kept running | "The underlying trigger was a GCP Persistent Disk I/O stall with an effectively unbounded timeout (provider default, to support live migration), which froze Consul, Patroni's HA loop, and other processes relying on the boot disk — the same disk used for cluster-membership state and logs." |
| 17 | GitLab | Corrective action: review failover candidate policy and document discovery path (INC-14518) | postmortem | 2026-10-01 | 2026-10-06 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23067 | The discovery path was not documented where responders needed it, which slowed diagnosis | "Patroni config details relevant to INC-14518 (consul.checks, remove_data_directory_on_*, maximum_lag_on_failover) are not documented alongside the discovery architecture, which slowed diagnosis during the incident." |
| 18 | GitLab | Corrective action: alert when a Patroni cluster is left paused (INC-14518) | postmortem | 2026-10-01 | 2026-10-06 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23070 | A paused control loop is an unalerted standing risk | "If a Patroni cluster is left paused, for example after maintenance ... there is no automatic failover for any reason and the watchdog is inactive. Nothing alerts on it today." |
| 19 | GitLab | Runbook: Primary unresponsive or no Leader | source | current master, retrieved 2026-10-06 | 2026-10-06 | https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/patroni/primary-unresponsive-assess-before-failover.md | The half-dead failure class, named and given a procedure: do not fail over first | "This runbook is for the rare half-dead primary: its Patroni or Consul agent is stuck (for example, on a stalled OS disk), so it loses the leader lock, while its Postgres keeps running and committing writes that the replicas no longer receive." Also: "the old primary is the most complete copy of the data", watchdog fires "85 s with our current settings", "Timebox: assessment ≤ 10 min." |
| 20 | GitLab | Runbooks MR 11682: Add Patroni leaderless-cluster and replica-streaming alerts | source | 2026-09, open at check time | 2026-10-06 | https://gitlab.com/gitlab-com/runbooks/-/merge_requests/11682 | The alerts that detect a leaderless cluster come from the control plane's own catalog, so they guard their own inputs with absent() | "PatroniNoHealthyPrimary |
| 21 | etcd | v3.5 data inconsistency postmortem | postmortem | 2022-04-20 | 2026-10-06 | https://github.com/etcd-io/etcd/blob/main/Documentation/postmortems/v3.5-data-inconsistency.md | The coordination store itself shipped a correctness bug for a year, and detection was unreliable by design | "Code refactor in v3.5.0 resulted in consistent index not being saved atomically. Independent crash could lead to committed transactions are not reflected on all the members." And: "Both checks however have a flaw, they depend on HashKV grpc method, which might fail causing the check to pass." |
| 22 | etcd | Avoiding Zombie Cluster Members When Upgrading to etcd v3.6 | postmortem | 2025-12-17 | 2026-10-06 | https://etcd.io/blog/2025/zombie_members_upgrade/ | Membership state in the coordinator diverged silently for years and resurfaced as removed members rejoining consensus | "This bug can cause the cluster to report \"zombie members\", which are etcd nodes that were removed from the database cluster some time ago, and are re-appearing and joining database consensus. The etcd cluster is then inoperable until these zombie members are removed." |
| 23 | etcd | Failure modes (op-guide, v3.5) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/etcd-io/website/blob/main/content/en/docs/v3.5/op-guide/failures.md | What the coordinator guarantees during its own failures: election pause for writes, no committed loss, lease extension | "During the leader election the cluster cannot process any writes." / "no committed writes are ever lost." / "The new leader extends timeouts automatically for all leases." / "When the majority members of the cluster fail, the etcd cluster fails and cannot accept more writes." |
| 24 | etcd | Why etcd (learning, v3.5) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/etcd-io/website/blob/main/content/en/docs/v3.5/learning/why.md | The coordinator chooses consistency over availability on purpose, so its clients inherit the availability problem | "These are systems that will never tolerate split-brain operation and are willing to sacrifice availability to achieve this end." |
| 25 | etcd | Tuning (op-guide, v3.5) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/etcd-io/website/blob/main/content/en/docs/v3.5/tuning.md | The detection constants: 100ms heartbeat, 1000ms election timeout, 50s upper limit | "By default, etcd uses a 100ms heartbeat interval." / "By default, etcd uses a 1000ms election timeout." / "The upper limit of election timeout is 50000ms (50s)". |
| 26 | Google (Mike Burrows) | The Chubby lock service for loosely-coupled distributed systems (OSDI '06; PDF mirror) | paper | 2006-11 | 2026-10-06 | https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf | Coarse-grained coordination is designed to be survivable; clients are told to expect losing it | "it is good for coarse-grained locks to survive lock server failures" and "Clients must be prepared to lose locks during network partitions, so the loss of locks on lock server fail-over introduces no new recovery paths." |
| 27 | Google (Mike Burrows) | Chubby (OSDI '06; PDF mirror) | paper | 2006-11 | 2026-10-06 | https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf | Election pauses are real and measured | "If a master fails, the other replicas run the election protocol when their master leases expire; a new master will typically be elected in a few seconds. For example, two recent elections took 6s and 4s, but we see values as high as 30s" |
| 28 | AWS (Brooker, Chen, Ping) | Millions of Tiny Databases (NSDI '20; PDF mirror) | paper | 2020-02 | 2026-10-06 | https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf | The 2011 EBS outage propagated because volume recovery depended on a control plane that shared a database with API traffic | "an incorrectly executed network configuration change triggered a condition which caused 13% of the EBS volumes in a single Availability Zone (AZ) to become unavailable. At that time, replication configuration was stored in the EBS control plane, sharing a database with API traffic." And quoting the postmortem: "The load caused a brown out of the EBS control plane and again affected EBS APIs across the Region." |
| 29 | AWS (Brooker, Chen, Ping) | Millions of Tiny Databases (NSDI '20; PDF mirror) | paper | 2020-02 | 2026-10-06 | https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf | Coordination load is highest exactly when the system is least able to serve it | "During normal operation, load consists of a low rate of calls caused by the background rate of EBS storage server failures, and creation of new cells for new volumes. During large-scale failures, however, load can increase considerably. This is an inherent risk of sometimes-coordinating systems like EBS: recovery load is not constant, and highest during bad network conditions." |
| 30 | AWS (Brooker, Chen, Ping) | Millions of Tiny Databases (NSDI '20; PDF mirror) | paper | 2020-02 | 2026-10-06 | https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf | The design goal: optimize availability of the coordinator conditional on the client needing it | "in terms of the availability of the volume (Av), and the instance (Ai), the control plane optimizes the conditional probability P(Av |
| 31 | Istio | Deployment models (docs) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/istio/istio.io/blob/master/content/en/docs/ops/deployment/deployment-models/index.md | Control plane scope is treated as outage scope, implying the mesh serves through control plane loss | "Improved availability: If a control plane becomes unavailable, the scope of the outage is limited to only workloads in clusters managed by that control plane." |
| 32 | Kubernetes | Kubernetes and Cluster Federation Control Plane Resilience | adr | 2015-12-14 | 2026-10-06 | https://github.com/kubernetes/design-proposals-archive/blob/main/multicluster/control-plane-resilience.md | Self-healing and availability of the control plane are separate properties, defined separately | "it's possible (but not desirable) to have high availability properties (e.g. multiple replicas) in the absence of self-healing properties (e.g. if a replica fails, nothing replaces it). Fairly obviously, given enough time, such systems typically become unavailable" |
| 33 | Patroni | Dynamic configuration (docs) | vendor | current master, retrieved 2026-10-06 | 2026-10-06 | https://github.com/patroni/patroni/blob/master/docs/dynamic_configuration.rst | The staleness budget constants: ttl 30s default with the invariant loop_wait + 2*retry_timeout <= ttl | "ttl: the TTL to acquire the leader lock (in seconds). Think of it as the length of time before initiation of the automatic failover process. Default value: 30" and "loop_wait + 2 * retry_timeout <= ttl" |
| 34 | Istio | Ambient multicluster performance (docs) | vendor | undated, retrieved 2026-10-06 | 2026-10-06 | https://github.com/istio/istio.io/blob/master/content/en/docs/ops/deployment/ambient-mc-perf/index.md | Measured control plane cost figure for scale planning | "Our multicluster control plane load test created 300 services with 4000 endpoints in each of 10 clusters ... The approximate control plane impact of adding a remote cluster at this scale was 1% of a CPU core, and 180 MB of memory." |
Tier mix: postmortem 5 rows / 4 artefacts, source 5, adr 6 rows / 4 artefacts, paper 5 rows / 2 artefacts, vendor 8, blog 3. 34 rows over 23 distinct artefacts, 4 hosts (github.com, raw.githubusercontent.com, gitlab.com, etcd.io).
Absences, stated: no conference talks (USENIX, KubeCon and YouTube unreachable from this session); no AWS Builders' Library (the canonical "static stability" texts); no Roblox 2021 Consul postmortem and no Cloudflare November 2020 etcd postmortem (both hosts unreachable) — the two most-cited public incidents in this failure class are represented here only by the same mechanisms appearing in reachable artefacts.