Coming back from cold: the restart is a harder load case than the crash
How production systems survive being switched on all at once: recovery load as a multiple of steady load, the warm-peer assumption inside warm-up machinery, and the dependency graph that must be acyclic at boot.
Reconstructs, from the Kubernetes project's own design records, the fixes HashiCorp shipped after October 2021, the Chubby and Physalia papers, Netflix and Meta operational docs, and GitLab's public incident tracker, why a full restart presents a demand curve steady-state capacity planning never saw. After reading, an architect can size the resync path (concurrent resyncers times response size times a measured 5x serialization multiplier), decide what a cold component may serve and when, and walk a boot-time dependency graph for the cycle that stops everything starting.
The industry's warm-up machinery states in its own documentation that it assumes something warm is still standing: Envoy's slow start calls itself 'not very effective' when all endpoints are new, mcrouter warms cold caches from a designated warm pool, and Eureka seeds a booting registry from a peer, so at total cold the engineered warm-up paths are inert and only jitter, admission control and disk remain.
What you get out of it
- Recovery work is a measured multiple of steady work: one client resync LIST costs the server about 5x the response size in transient memory, and sixteen synchronized resyncers killed the Kubernetes project's test apiserver (KEP-3157, 2022).
- Dispersal must live in the client contract: gRPC's backoff protocol requires that connection backoffs started at the same time disperse, because a server restart is exactly that case and the balancer's slow start is inert when every endpoint is cold.
- A component that boots empty should distrust its own emptiness: Eureka quarantines an unseeded registry for a shipped default of five minutes rather than serve partial truth, and Chubby's failed-over master admits KeepAlives before any session operation.
- The boot-time dependency graph is not the runtime graph, and it must be acyclic: kubeadm's self-hosted control plane could not survive a reboot because the pivot deleted its static bootstrap rung, and the kube-proxy readiness probe was rejected because probing the fixer restarts it into a herd.
- Recovery schedules the next incident unless you drain it deliberately: GitLab's cold Redis cache and its re-enabled gc counter each turned one day's recovery into the next day's outage, nine years apart.
Scope
Why this, now. The best-documented total-restart incidents cluster in the last five years and the Kubernetes streaming-resync machinery that answers them only became default recently, so the gap between what teams run and what the record teaches is at its widest.
What it does not cover. Datacenter power sequencing, region failover, steady-state cache stampedes, retry-storm dynamics and control-plane loss while serving, which are covered by sibling guides; the AWS, Meta, Roblox and Cloudflare restart postmortems hosted on domains unreachable from this session are represented only through the reachable primary record and are named as absences in the ledger.
Other field guides
The way back is also a deploy
Reconstructs rollback as it actually works in production: a forward deploy of an old artifact, underwritten by a compatibility contract fixed weeks b…
28 sources · 25 organisations · 6 postmortemsWhere resilience policy lives: Netflix, 2016 to 2026
Between 2016 and 2026 almost every Netflix library that decided something about the network was retired, while the libraries that decide something ab…
22 sources · 4 organisations · 4 postmortemsDeciding who gets told no
A field guide to request-level admission control, built from GitLab's public incident tracker and six years of its rate-limiting change record, the E…
26 sources · 13 organisations · 4 postmortems