Every source behind this page, graded. This session’s network reached
code hosts only, so the wall is repository artefacts, GitLab’s public tracker, and two
paper mirrors; there are no talks, and the engineering-blog layer is thin by constraint
rather than choice.
Postmortem
GitLab2026-09
INC-14518 corrective action: Patroni fencing and self-shutdown
The public record of the half-dead primary: a GCP disk stall froze Consul and Patroni,
the node lost its lock, and Postgres kept committing for about 33 minutes because the
watchdog was never configured and /dev/watchdog did not exist.
Carry forwardDemotion is a decision; fencing is a mechanism. Only the mechanism counts when the agents are frozen.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/23046
Postmortem
GitLab2026-10
INC-14518 corrective actions: discovery path and paused clusters
Two more public follow-ups: the Patroni/Consul/pgbouncer discovery path was undocumented,
which slowed diagnosis; and a paused Patroni cluster (no failover, no watchdog) raised no
alert at all.
Carry forwardA disabled control loop is a standing incident with no alarm. Alert on the loop’s state, not just its outputs.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/23070
Postmortem
etcd maintainers2022-04
v3.5 data inconsistency postmortem
Committed to the repository itself: a refactor broke the atomicity of the consistent
index, members could diverge after crashes, and the corruption check depends on a call
that can fail in a way that passes the check.
Carry forwardThe coordinator is software. Run its integrity checks and rehearse restoring it; do not model it as a constant.
github.com/etcd-io/etcd/.../v3.5-data-inconsistency.md
Postmortem
etcd maintainers2025-12
Avoiding zombie cluster members when upgrading to v3.6
Years-old divergence between etcd’s two membership stores resurfaced as removed
members rejoining consensus on upgrade, leaving clusters inoperable; the triggers are
old disaster-recovery actions.
Carry forwardEvery force-recovery of the coordinator leaves a scar; inventory them before the next coordinator upgrade.
etcd.io/blog/2025/zombie_members_upgrade
Decision record
Kubernetespre-2018
Design principles: continue on last instructions
The founding architecture principles include the whole pattern in one line: components
should continue to do what they were last told in the absence of new instructions, with
partition and component outage named as the motivating cases.
Carry forwardMake staleness tolerance a stated principle, so every new controller inherits it instead of re-deciding it.
github.com/kubernetes/design-proposals-archive/.../principles.md
Decision record
Envoychecked 2026
xDS protocol: last-known-good, and the TTL exception
The protocol document states the fail-static default and then the exception that proves
the rule: fault-injection config left behind by a dead management server is dangerous, so
resources can carry a TTL that removes them on contact loss.
Carry forwardClassify every config kind as safe-stale or dangerous-stale, and attach expiry only to the second.
github.com/envoyproxy/envoy/.../xds_protocol.rst
Decision record
Patronichecked 2026
DCS failsafe mode: the rejected quorum, in writing
The design document for what a primary may do when the DCS is gone: demote by default,
survive only on unanimous member acknowledgement, and an FAQ that rejects quorum because
the DCS and the database cluster can hold different views of the same partition.
Carry forwardWhen two systems each have a quorum, their intersection is not one; unanimity is the only cheap safe check.
github.com/patroni/patroni/blob/master/docs/dcs_failsafe_mode.rst
Decision record
Kubernetes2015-12
Control plane resilience design (Hoole, Danese, Santa-Barbara)
Separates self-healing from high availability and notes that replicas without
replacement “fairly obviously” drift to unavailability over time; the control
plane needs both properties, independently engineered.
Carry forwardAn HA coordinator that nothing repairs is a slow outage; budget the repair loop, not just the replica count.
github.com/kubernetes/design-proposals-archive/.../control-plane-resilience.md
Source code
Kubernetesmaster, 2026
node_lifecycle_controller.go: zone states and master disruption mode
The disbelief gate in executable form: ComputeZoneState flips a zone to
partialDisruption at 55% NotReady, and when every zone is fully disrupted the controller
marks all nodes reachable and swaps every eviction limiter to zero.
Carry forwardRate-limit destructive reactions by the fraction of the fleet affected, with a hard stop at “everyone looks dead.”
github.com/kubernetes/kubernetes/.../node_lifecycle_controller.go
Source code
Kubernetesmaster, 2026
kube-controller-manager node lifecycle flags
The shipped numbers with their own documentation: eviction at 0.1 nodes/s when healthy,
0.01 when a zone is unhealthy, overridden to zero below 50 nodes, threshold 0.55 with a
minimum of three NotReady nodes.
Carry forwardSmall clusters get no secondary eviction at all; upstream decided that for you, and it is probably right.
github.com/kubernetes/kubernetes/.../options/nodelifecyclecontroller.go
Source code
Netflixmaster, 2026
Eureka: PeerAwareInstanceRegistryImpl and server config
Lease expiry runs only while renewals exceed the threshold
(renewalPercentThreshold 0.85); the registry would rather serve dead entries than evict
live ones on a broken signal.
Carry forwardExpiry driven by heartbeats must be gated on the heartbeat channel itself looking healthy.
github.com/Netflix/eureka/.../PeerAwareInstanceRegistryImpl.java
Source code
GitLab2026-10
Runbook: primary unresponsive, recover without losing data
Operational doctrine written after INC-14518: do not fail over first, because the
half-dead primary is the most complete copy of the data; with their settings the watchdog
fires 85 seconds after the last successful lock renewal; assessment is timeboxed to 10
minutes.
Carry forwardWrite the “half-dead” branch into the runbook before the incident; it is the branch responders get wrong under pressure.
gitlab.com/gitlab-com/runbooks/.../primary-unresponsive-assess-before-failover.md
Source code
GitLab2026-10, open
Runbooks MR 11682: leaderless-cluster alerts that guard their own inputs
The proposed alerts detect a leaderless Patroni cluster from Consul’s catalog, and
a third, non-paging alert fires when the inputs of the first two go absent, because an
alert fed by the control plane disappears with it.
Carry forwardAny alert whose data flows through the coordinator needs an absent() guard, or the outage silences its own alarm.
gitlab.com/gitlab-com/runbooks/-/merge_requests/11682
Paper
Google (Burrows)2006-11
The Chubby lock service (OSDI ’06; PDF mirror)
The origin text for coarse-grained coordination that survives its own failures: five
replicas per cell, elections measured at 4 to 6 seconds with a 30-second tail, and the
client contract that locks can vanish in partitions without creating new recovery paths.
Carry forwardDesign clients so losing coordination is a slow path they already have, not a new code path discovered in production.
raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf
Paper
AWS (Brooker, Chen, Ping)2020-02
Millions of Tiny Databases (NSDI ’20; PDF mirror)
Physalia: consensus for EBS re-sharded into per-volume cells after the 2011 outage,
designed to maximise the probability the coordinator is up conditional on a client
needing it, with the warning that recovery load peaks during bad network conditions.
Carry forwardMeasure your coordinator’s availability conditioned on failure events, not averaged over calendar time.
raw.githubusercontent.com/arpit20adlakha/.../Millions of Tiny Databases.pdf
Eng blog
Netflixchecked 2026
Eureka wiki: self-preservation and peer-to-peer behaviour
Netflix-authored operational doctrine: stop evicting when more than 15% of the registry
looks dead at once, hold a cold server for five minutes rather than serve a partial
registry, and tell clients plainly that they must tolerate stale entries.
Carry forwardA registry that can be emptied by a network event will be; prefer stale-and-full to fresh-and-empty.
github.com/Netflix/eureka/wiki/Server-Self-Preservation-Mode
Eng blog
Netflixchecked 2026
Eureka at a glance: the client cache as the availability story
States the architecture bet directly: discovery data cached in every client means
applications resilient to the outage of the registry itself, contrasted explicitly with
proxy-based load balancing where the balancer’s outage is yours.
Carry forwardMoving a dependency out of the request path and into a cached instruction is the cheapest nine you will ever buy.
github.com/Netflix/eureka/wiki/Eureka-at-a-glance
Vendor docs
Envoychecked 2026
Panic threshold, with the fail-open/fail-closed switch
Official docs for the 50% default, the two panic behaviours, and the stated condition
for choosing fail-closed: upstreams that fail all-or-nothing, where partial traffic only
deepens the failure.
Carry forwardDecide per upstream whether some success beats fast failure; the default assumes it does.
github.com/envoyproxy/envoy/.../panic_threshold.rst
Vendor docs
etcdv3.5, checked 2026
Failure modes, tuning, and why etcd
The coordinator’s own contract: writes pause during elections but committed writes
survive; leases are auto-extended across elections; majority loss stops writes entirely;
and the explicit statement that availability is sacrificed for split-brain safety.
Carry forwardRead your coordinator’s failure-modes page as an SLA; everything it declines to promise is your staleness budget.
github.com/etcd-io/website/.../op-guide/failures.md
Vendor docs
Istiochecked 2026
Deployment models and control-plane cost
Control-plane scope treated as outage scope (more control planes, smaller blast
radius), plus a vendor load test pricing a remote cluster at roughly 1% of a CPU core and
180 MB of control-plane memory at 300 services and 4,000 endpoints.
Carry forwardControl planes are cheap enough to multiply; blast radius, not cost, should set their count.
github.com/istio/istio.io/.../deployment-models/index.md
Vendor docs
Patronichecked 2026
Dynamic configuration: the lease arithmetic
TTL 30 seconds by default, minimum 20, with the invariant loop_wait + 2×retry_timeout
≤ ttl that bounds how many verification attempts fit inside one lease; member slots
retained 30 minutes after a member vanishes.
Carry forwardYour failover time is a sum you can compute from config; compute it before an incident does.
github.com/patroni/patroni/blob/master/docs/dynamic_configuration.rst