websocketsServices error rate at 16.14% on frontend main
A five-minute burst of connection resets became a reconnect herd: 144,000 sessions, autoscaling from ~100 to 335 pods. The trigger was never identified; the cost ratio was measured anyway.
How production systems survive being switched on all at once, reconstructed from the Kubernetes project's own design records, the fixes HashiCorp shipped after October 2021, Google's Chubby and EBS papers, Netflix and Meta operational docs, and GitLab's public incident tracker. After reading, you can size recovery load as a multiple of steady load, spot the warm-peer assumption hiding in your warm-up machinery, and walk your dependency graph for the cycle that stops the whole thing booting.
The problem, stated without naming a technology: after a total stop, a system must rebuild in minutes the state it normally accumulates over hours, while every client it has ever had asks for everything at the same time.
Steady state is the easy case. A fleet in steady state holds an enormous amount of accumulated context that nobody provisioned explicitly: caches at their hit rate, watch connections established, session state in server memory, connection pools filled, registries populated by heartbeats. The load the system was sized for assumes all of it. A full restart deletes all of it at once, and then presents a demand curve that steady-state capacity planning never saw: every client reconnecting in the same second, every cache miss going to the backing store, every watcher re-listing the world from scratch, plus the backlog of real work that accumulated during the outage.
The surprise in this material, and the finding this guide leads with, is that the industry's warm-up machinery is explicitly not built for that case. The three best documented warm-up mechanisms all state, in their own documentation, that they assume something warm is still standing. Envoy's slow start ramps traffic into a new endpoint by shifting load onto warm ones, and its docs say directly that when all endpoints are new "slow start is not very effective". Meta's mcrouter warms a cold cache by copying values from a designated warm pool. Netflix's Eureka primes a booting registry from a neighbouring node. I am calling this the warm-peer assumption; the sources do not share a name for it. Partial cold is an engineered, well-trodden path. Total cold is a different regime in which that path is inert, and the mechanisms that still work are the unglamorous ones: jittered backoff, admission control, and a bootstrap rung that reads from disk instead of from a peer.
Scope. This guide covers the restart-from-zero of a service fleet and the coordination layer under it: what arrives at T+0, which mechanisms disperse it, what a component may serve while cold, and what the dependency graph must look like for the bottom layer to boot at all. It deliberately does not cover datacenter power sequencing, region failover, or three adjacent failure loops this practice has already dug: retry storms that sustain an overload, the steady-state cache stampede, and data planes serving while their control plane is down. The canonical restart incidents at AWS, Meta and Roblox are hosted on domains unreachable from this session's network; their mechanics enter here through the reachable primary record (the fixes HashiCorp shipped, the Kubernetes project's design records) and through the sibling guide on name-resolution failure, which cites them directly. The evidence ledger states this constraint precisely.
The common shape across Chubby, Eureka, the Kubernetes API machinery, Consul and mcrouter: dispersal at the client, admission at the door, a cold phase that refuses to improvise, and a bootstrap rung that depends on nothing alive.
Dispersal is a contract, not a courtesy. gRPC's connection backoff protocol fixes the constants every compliant client ships: initial backoff 1 second, multiplier 1.6, cap 120 seconds, jitter ±20%, and then adds a sentence that is really a law about restarts: "Alternate implementations must ensure that connection backoffs started at the same time disperse". Started at the same time is not an edge case; it is the definition of a server restart. Consul was retrofitting the same idea in February 2015, in its 0.5.0 changelog: "Random stagger for agent checks to prevent thundering herd". The transfer for your own design: the dispersal has to live in the client contract, because by the time the herd arrives the server has no good options left, and a load balancer cannot ramp traffic between endpoints when every endpoint is equally cold.
The door has a priority order before it has a queue. Kubernetes' API Priority and Fairness KEP opens its motivation with the overload scenario in which the requests that keep the cluster alive lose: "Self-Maintenance crowded out... In an overload scenario today there is no assurance of priority for these self-maintenance requests". Chubby's failover protocol is the same idea expressed as a ladder: a new master first answers only KeepAlives, then session operations, then everything else, and it rebuilds its session table "partly by obtaining state from clients". The clients are not only the problem at T+0; in Chubby's design they are also the warm source of last resort before disk.
The cold phase refuses to improvise. A component that boots empty has a choice
about what its emptiness means. Eureka's answer is the most explicit in the public record:
a registry that cannot sync from any peer waits five minutes, a shipped constant
(waitTimeInMsWhenSyncEmpty, 1000 * 60 * 5)
while clients re-register, and the wiki gives the reason: the server "tries
hard not to provide partial information to the clients there by skewing traffic only to a
group of instances and causing capacity issues". An empty registry served as truth
would tell every client that every service is gone, or worse, that three instances are the
whole fleet. The kube-apiserver has the same shape at a different layer: its watch cache
initialises from the store "with
empty change history", so returning watchers cannot resume and fall back to full
re-lists, which is why KEP-1904 teaches the cache to initialise from a compacted revision
instead. Both are instances of one rule: at boot, distrust your own emptiness.
The bottom rung reads from disk and asks permission of nobody. etcd, the coordinator under Kubernetes and under half the systems cited here, states its own floor plainly: below quorum it "disastrously fails, irrevocably losing quorum", and the way back is a human with a snapshot file. The restore deliberately refuses continuity: restoring "overwrites some snapshot metadata (specifically, the member ID and cluster ID); the member loses its former identity", so a restored node cannot half-join the wreck of the old cluster. Kubernetes learned the corresponding lesson about layers above: kubeadm's experiment in running the control plane on the control plane (self-hosting) shipped with the caveat that a self-hosted cluster "cannot recover from a reboot of the control-plane node without manual intervention", because the pivot's final step deletes the static-pod manifests that were the bootstrap rung. The feature stayed experimental and the static pods stayed, etcd included. State of this claim: the caveat and the mechanism are reported; the inference that the caveat is why self-hosting never became the default is mine, and the design record is consistent with it.
Each of these was argued in public, with a rejected side and a stated reason. The flips-when column is the part to carry into your own review.
| Decision | Chosen | Rejected | Because | Evidence |
|---|---|---|---|---|
| Serve while cold? | Quarantine until synced | Serve immediately | Partial registry skews all traffic onto a few instances | Eureka wiki |
| Dispersal location | Client contract (backoff + jitter) | Balancer ramp | Slow start is inert when every endpoint is cold | Envoy docs, gRPC protocol |
| Resync shape | Resume / stream | Full LIST per reconnect | O(5x) transient memory; 16 informers killed the server | KEP-3157 |
| Platform on itself | Static bootstrap rung kept | Full self-hosting | "Cannot recover from a reboot... without manual intervention" | kubeadm docs |
| Probe the foundation? | Metrics and logs | Readiness probe on kube-proxy | A probe-driven restart of the fixer is itself the herd | PR #111857, closed unmerged |
| Raft log durability bookkeeping | bbolt + NoFreelistSync option | Legacy boltdb defaults | Freelist sync cost surfaced at scale after October 2021 | Consul PR #11720 |
The fifth row deserves a sentence, because it is the rare rejected pull request whose rejection reasoning is the whole topic in miniature. A contributor proposed a readiness probe for kube-proxy, the component that programs every node's service routing. The reviewer's argument against, quoted in full in the ledger, is that under CPU pressure the health endpoint flaps, the probe then restarts kube-proxy everywhere at once, and "kube-proxy trying to bootstrap from scratch and all the pods on the nodes reconnecting to the Services we just broke" takes the cluster down. The PR closed unmerged in August 2022. Health-checking is not free at the foundation layer: a probe is a restart trigger, and a restart of the foundation is a cold start of everything above it.
Four failure classes recur: the synchronized client, recovery amplification, the cold-cache aftershock, and the fixer trapped inside the failure. Each card ends with the rule to carry forward.
One class is conspicuously thin in the reachable record: no public account measures a full cold start end-to-end, from power-on to steady-state hit rates, with enough numbers to plan against. The component-level evidence above is the closest thing. Treat your own first full-restart drill as the measurement that does not exist publicly, and instrument it accordingly.
Everything quantitative in the corpus, with context and dates. Measured unless marked otherwise.
| Metric | Value | At | Context | As of | Source |
|---|---|---|---|---|---|
| Reconnect backoff constants | 1s · ×1.6 · cap 120s · jitter ±20% | gRPC | Mandated for every compliant client | 2017+ | protocol doc |
| Transient memory per resync LIST | ~5× response size | Kubernetes | Measured on the serving path, etcd-backed LIST | 2022 | KEP-3157 |
| Synchronized clients to kill one server | 16 informers | Kubernetes | 400 secrets of 1 MB each; synthetic reproduction | 2022 | KEP-3157 |
| Cache-priming data moved | ~400 GB in ~25 min | Kubernetes | Same test; network cost of re-listing | 2022 | KEP-3157 |
| Worked resync sizing, before and after streaming | 2.5 TB → 2 GB | Kubernetes | 512 watchers × 400 MB dataset (derived in the KEP: 512×500×2MB×5) | 2022 | KEP-3157 |
| Reconnect burst fleet ratio | ~100 → 335 pods | GitLab | 144,000 sessions reset over 5 minutes | 2026-10 | incident #23063 |
| Cold-cache stopgap capacity | 7 → 11 web nodes | GitLab | Day-after outage served on an empty Redis | 2017-09 | incident #250 |
| Empty-boot quarantine | 5 min | Netflix Eureka | Registry boots with no warm peer; shipped default | current master | source |
| Self-preservation threshold | 85% renewals / 15 min | Netflix Eureka | Below it, the registry stops expiring instances | wiki, undated | wiki |
| Failover grace period / lease extension | 45s / 12s | Google Chubby | Clients ride out a master failover without losing sessions | 2006 | OSDI '06 |
| Clients per coordinator | 90,000 | Google Chubby | "clients can overwhelm the master by a huge margin" | 2006 | OSDI '06 |
| Subscription fan-out benchmark | 70,000 subs / 128 cores | HashiCorp Consul | Event-buffer contention fix validation | 2022-01 | PR #12080 |
| Slow-start minimum weight | 10% | Envoy | Floor for a cold endpoint's share during the ramp window | current docs | docs |
| Recovery blast radius | 2.5% YouTube views (1h); 30% GCS traffic | Multi-region config error; slow restore through congested fabric | 2019-06 | blog |
The Kubernetes figures are the project's own synthetic measurements, not a customer incident; the 2.5 TB row is derived arithmetic stated in the KEP, not an observation. The Chubby constants are two decades old and prove the shape, not current values. Nobody in this corpus publishes an end-to-end cold-start duration for a full production stack; the Cloudflare November 2023 account reportedly does, but its host was unreachable from this session and it is deliberately not cited here.
Every source behind this page, graded. Four postmortems, four source records, five design records, three operator docs, two wikis, two papers. No talks: the venues that host them were unreachable from this session's network, and the ledger says so rather than padding the tier.
A five-minute burst of connection resets became a reconnect herd: 144,000 sessions, autoscaling from ~100 to 335 pods. The trigger was never identified; the cost ratio was measured anyway.
A guard disabled during one incident let a trigger counter accumulate fleet-wide; the next deploy re-enabled the check and scheduled all the deferred work at once.
The day after an outage, ordinary traffic on a cold Redis cache drove worker timeouts site-wide; capacity (7 to 11 web nodes) was the stopgap while the cache warmed.
Detection took seconds; restoration took hours, because the congestion being fixed also carried the engineers' tooling. YouTube lost 2.5% of views for an hour, GCS 30% of traffic.
Names the restart, not the crash, as the recurring trigger: a rebooted server has empty watch history, so every returning client is forced into a full relist.
Measures the resync path: ~5x response size of transient memory per LIST, a server lost at 16 concurrent informers, and large-cluster recovery named as "a very expensive storm of LISTs". The fix streams initial state: 2.5 TB to 2 GB in the worked example.
The motivation list is a taxonomy of overload at the door, led by "Self-Maintenance crowded out": without explicit priority, the requests that keep the system alive lose to the herd.
The client-side half of restart survival as a compliance document: exponential backoff, stated constants, and the requirement that simultaneous backoffs disperse.
The procedure that made the control plane self-hosted ends by deleting the static-pod manifests, which is precisely the bootstrap rung a reboot would need.
Rejected because a flapping health endpoint would restart the routing layer everywhere at once: "kube-proxy trying to bootstrap from scratch and all the pods on the nodes reconnecting... will bring the whole cluster down".
"Random stagger for agent checks to prevent thundering herd": the dispersal rule, shipped eleven years ago, in the changelog of the system at the centre of the best-known total-restart incident.
Weeks after October 2021, Consul's raft log store moved to bbolt and exposed freelist-sync control, with new write-capacity metrics. The reachable artefact of the unreachable postmortem.
At 70k subscriptions on 128 cores, publishing one event woke every subscriber; "the majority of scheduled goroutines have no real work to do". Contention lived in the kernel's futex table, not the application locks.
"A self-hosted cluster cannot recover from a reboot of the control-plane node without manual intervention", stated as a product caveat on the feature that removed the static bootstrap rung.
On boot the registry syncs from a neighbour; failing that it quarantines itself,
because partial information "skew[s] traffic only to a group of instances and caus[es]
capacity issues". The quarantine is a shipped constant,
waitTimeInMsWhenSyncEmpty,
defaulting to 1000 × 60 × 5.
The operator's warm-up: misses on the cold pool are filled from a designated warm pool, protecting the backing store. The design presumes the warm pool survived.
Weight-ramps new endpoints over a window, minimum 10% share, and states its own limit: with every endpoint new, "slow start is not very effective".
Below quorum the cluster "disastrously fails"; recovery is a human with a snapshot,
and the restored member deliberately loses its identity so it cannot rejoin the old
cluster. --force-new-cluster exists and is "strongly discouraged".
The failover protocol is an admission ladder: rebuild sessions from disk and from the clients themselves, serve KeepAlives before operations. 90,000 clients per master; grace period 45s.
The recovery coordinator's load profile, stated exactly: near-idle in normal operation, then a latency-critical "burst of work" when a large-scale failure takes many servers offline at once, "most critical at the most challenging time".
Six rungs. The line from toy to real is crossed at rung four, when you start killing everything at once on purpose.
One server, 500 client processes holding long-lived connections, no backoff logic. Kill the server for ten seconds, bring it back, graph connections per second and server CPU on reattach.
Done when: you can show the reconnect spike as a multiple of steady connection rate. Teaches: the restart is a load event, generated by your own clients.
Implement the gRPC backoff algorithm (1s initial, 1.6 multiplier, 120s cap, 20% jitter) in your client. Repeat the kill. Then remove only the jitter and watch the synchronized waves return.
Done when: reattach load spreads across the backoff window, and you have seen why jitter, not backoff, does the spreading. Teaches: dispersal is the client contract's job.
Put a cache in front of a deliberately slow store (add 50ms per read). Restart the cache empty under steady load; measure the amplification at the store and the time back to steady hit rate. Add a warm-peer route of the mcrouter kind and measure again.
Done when: you can state the store's cold-start amplification factor and the warm route's effect on it. Teaches: the aftershock: the outage is over, the miss storm is not.
Now restart server, cache and clients together. The warm route has nothing warm to read from; watch it go inert. Make the server refuse full-resync requests beyond a concurrency cap (admission control) and quarantine itself for N seconds while clients re-register, Eureka-style.
Done when: total cold start completes without the store or the server falling over, slower but monotonically. Teaches: the warm-peer assumption, by taking the peer away.
Give health checks, registrations and leader heartbeats their own priority class at the server's door, with reserved concurrency, per KEP-1040's motivation. Re-run rung four and confirm the system can see itself while overloaded.
Done when: heartbeats hold their latency through the resync storm. Teaches: fairness is not enough; survival traffic needs a reservation.
Productionising: draw the dependency graph of your real stack with one question per edge, "does this edge exist at boot?". Find the cycle (there is usually one, often through DNS, secrets, or the deploy system). Break it with a static rung, write the black-start runbook, then run a scheduled full-restart drill and measure power-on to steady-state hit rates.
Done when: the drill produces the end-to-end cold-start number nobody publishes. Teaches: the graph that must be acyclic is the boot-time graph, which is not the runtime graph.
The queries and moves that found this material, in the forms that worked. The wiki-raw and changelog tricks travel to any project.
kube-apiserver restart LIST storm OOM "thundering herd" relist"reconnect storm" OR "thundering herd" site:gitlab.com gl-infra production"cold start" "black start" postmortem restart dependency cycle bootstrap"waitTimeInMsWhenSyncEmpty"repo:kubernetes/kubernetes "thundering herd" is:pr is:closed is:unmergedpath:keps "after kube-apiserver reboot" OR "storm of LISTs"grpc "connection backoff" protocol "must ensure" dispersegrep -n -i "thundering\|stagger\|freelist\|backpressure" CHANGELOG.mdconsul changelog "boltdb" freelist bbolt 1.11https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=cold+cachehttps://raw.githubusercontent.com/wiki/<org>/<repo>/<Page>.mdsite:github.com wiki "cold cache" warm up mcrouter