COLD START  / field guide
Practitioner field guide · 2026-10-07

Coming back from cold: the restart is a harder load case than the crash

How production systems survive being switched on all at once, reconstructed from the Kubernetes project's own design records, the fixes HashiCorp shipped after October 2021, Google's Chubby and EBS papers, Netflix and Meta operational docs, and GitLab's public incident tracker. After reading, you can size recovery load as a multiple of steady load, spot the warm-peer assumption hiding in your warm-up machinery, and walk your dependency graph for the cycle that stops the whole thing booting.

21 primary sources 9 organisations 4 published incidents Evidence through October 2026 Read: 26 min
01

The territory

The problem, stated without naming a technology: after a total stop, a system must rebuild in minutes the state it normally accumulates over hours, while every client it has ever had asks for everything at the same time.

5×
Temporary server memory per client resync request, measured by the Kubernetes project
16
Synchronized cache-priming clients that killed an API server in the project's own test
3.3×
Fleet growth GitLab's autoscaler needed to absorb one five-minute reconnect burst
5 min
How long a Eureka registry that boots empty serves nothing rather than serving lies

Steady state is the easy case. A fleet in steady state holds an enormous amount of accumulated context that nobody provisioned explicitly: caches at their hit rate, watch connections established, session state in server memory, connection pools filled, registries populated by heartbeats. The load the system was sized for assumes all of it. A full restart deletes all of it at once, and then presents a demand curve that steady-state capacity planning never saw: every client reconnecting in the same second, every cache miss going to the backing store, every watcher re-listing the world from scratch, plus the backlog of real work that accumulated during the outage.

The surprise in this material, and the finding this guide leads with, is that the industry's warm-up machinery is explicitly not built for that case. The three best documented warm-up mechanisms all state, in their own documentation, that they assume something warm is still standing. Envoy's slow start ramps traffic into a new endpoint by shifting load onto warm ones, and its docs say directly that when all endpoints are new "slow start is not very effective". Meta's mcrouter warms a cold cache by copying values from a designated warm pool. Netflix's Eureka primes a booting registry from a neighbouring node. I am calling this the warm-peer assumption; the sources do not share a name for it. Partial cold is an engineered, well-trodden path. Total cold is a different regime in which that path is inert, and the mechanisms that still work are the unglamorous ones: jittered backoff, admission control, and a bootstrap rung that reads from disk instead of from a peer.

Figure 1 · Two load regimes, one capacity plan

T+0 of a restart: demand nobody provisioned

Steady state: demand the system was sized for

restart deletes the
accumulated state

Reads mostly absorbed
by warm caches

Server at planned
utilisation

Watches and sessions
already established

Every client
reconnects at once

Server with cold cache,
empty history, no sessions

Every cache miss
hits the backing store

Every watcher re-lists
the full dataset

Deferred work arrives
as one batch

T+0 of a restart: demand nobody provisioned

Steady state: demand the system was sized for

restart deletes the
accumulated state

Reads mostly absorbed
by warm caches

Server at planned
utilisation

Watches and sessions
already established

Every client
reconnects at once

Server with cold cache,
empty history, no sessions

Every cache miss
hits the backing store

Every watcher re-lists
the full dataset

Deferred work arrives
as one batch

Four demand sources that are quiet in steady state arrive together at T+0 of a restart, and they land on a server that has also lost its own accelerating state. Reconstructed from KEP-1904, GitLab #23063 and Millions of Tiny Databases.
Diagram source

Scope. This guide covers the restart-from-zero of a service fleet and the coordination layer under it: what arrives at T+0, which mechanisms disperse it, what a component may serve while cold, and what the dependency graph must look like for the bottom layer to boot at all. It deliberately does not cover datacenter power sequencing, region failover, or three adjacent failure loops this practice has already dug: retry storms that sustain an overload, the steady-state cache stampede, and data planes serving while their control plane is down. The canonical restart incidents at AWS, Meta and Roblox are hosted on domains unreachable from this session's network; their mechanics enter here through the reachable primary record (the fixes HashiCorp shipped, the Kubernetes project's design records) and through the sibling guide on name-resolution failure, which cites them directly. The evidence ledger states this constraint precisely.

02

How restartable systems are actually built

The common shape across Chubby, Eureka, the Kubernetes API machinery, Consul and mcrouter: dispersal at the client, admission at the door, a cold phase that refuses to improvise, and a bootstrap rung that depends on nothing alive.

Figure 2 · Reference architecture for surviving your own restart

Warm sources, in preference order

Restarted service

Client population (the herd)

reconnects, dispersed
over the backoff window

sync before serving

must depend on
nothing alive

Clients with mandated
exponential backoff + jitter

Admission layer:
priority + flow control,
self-maintenance reserved

Cold gate: refuse or limit
service until synced

Serving layer
(cache, sessions, watches)

1. A warm peer
(registry copy, warm cache pool)
2. The clients themselves
(state handed back on reconnect)
3. Durable local state
(snapshot, static manifest)

Bootstrap rung:
files on disk, static config

Warm sources, in preference order

Restarted service

Client population (the herd)

reconnects, dispersed
over the backoff window

sync before serving

must depend on
nothing alive

Clients with mandated
exponential backoff + jitter

Admission layer:
priority + flow control,
self-maintenance reserved

Cold gate: refuse or limit
service until synced

Serving layer
(cache, sessions, watches)

1. A warm peer
(registry copy, warm cache pool)
2. The clients themselves
(state handed back on reconnect)
3. Durable local state
(snapshot, static manifest)

Bootstrap rung:
files on disk, static config

Every layer exists in at least two of the cited systems. The right-hand bootstrap path is the one that must not depend on anything in the left column. Reconstructed from Chubby, Eureka, KEP-1040 and etcd's recovery guide.
Diagram source

Dispersal is a contract, not a courtesy. gRPC's connection backoff protocol fixes the constants every compliant client ships: initial backoff 1 second, multiplier 1.6, cap 120 seconds, jitter ±20%, and then adds a sentence that is really a law about restarts: "Alternate implementations must ensure that connection backoffs started at the same time disperse". Started at the same time is not an edge case; it is the definition of a server restart. Consul was retrofitting the same idea in February 2015, in its 0.5.0 changelog: "Random stagger for agent checks to prevent thundering herd". The transfer for your own design: the dispersal has to live in the client contract, because by the time the herd arrives the server has no good options left, and a load balancer cannot ramp traffic between endpoints when every endpoint is equally cold.

The door has a priority order before it has a queue. Kubernetes' API Priority and Fairness KEP opens its motivation with the overload scenario in which the requests that keep the cluster alive lose: "Self-Maintenance crowded out... In an overload scenario today there is no assurance of priority for these self-maintenance requests". Chubby's failover protocol is the same idea expressed as a ladder: a new master first answers only KeepAlives, then session operations, then everything else, and it rebuilds its session table "partly by obtaining state from clients". The clients are not only the problem at T+0; in Chubby's design they are also the warm source of last resort before disk.

The cold phase refuses to improvise. A component that boots empty has a choice about what its emptiness means. Eureka's answer is the most explicit in the public record: a registry that cannot sync from any peer waits five minutes, a shipped constant (waitTimeInMsWhenSyncEmpty, 1000 * 60 * 5) while clients re-register, and the wiki gives the reason: the server "tries hard not to provide partial information to the clients there by skewing traffic only to a group of instances and causing capacity issues". An empty registry served as truth would tell every client that every service is gone, or worse, that three instances are the whole fleet. The kube-apiserver has the same shape at a different layer: its watch cache initialises from the store "with empty change history", so returning watchers cannot resume and fall back to full re-lists, which is why KEP-1904 teaches the cache to initialise from a compacted revision instead. Both are instances of one rule: at boot, distrust your own emptiness.

The bottom rung reads from disk and asks permission of nobody. etcd, the coordinator under Kubernetes and under half the systems cited here, states its own floor plainly: below quorum it "disastrously fails, irrevocably losing quorum", and the way back is a human with a snapshot file. The restore deliberately refuses continuity: restoring "overwrites some snapshot metadata (specifically, the member ID and cluster ID); the member loses its former identity", so a restored node cannot half-join the wreck of the old cluster. Kubernetes learned the corresponding lesson about layers above: kubeadm's experiment in running the control plane on the control plane (self-hosting) shipped with the caveat that a self-hosted cluster "cannot recover from a reboot of the control-plane node without manual intervention", because the pivot's final step deletes the static-pod manifests that were the bootstrap rung. The feature stayed experimental and the static pods stayed, etcd included. State of this claim: the caveat and the mechanism are reported; the inference that the caveat is why self-hosting never became the default is mine, and the design record is consistent with it.

Figure 3 · The boot ladder a cold coordinator climbs

process starts,
memory empty

read disk snapshot /
ask a warm peer

no peer answered

wait out the quiet period
(Eureka default 5 min)
while clients re-register

state rebuilt
(conservative approximation)

heartbeats / KeepAlives only,
then session ops (Chubby steps 4-8)

Cold

Syncing

Quarantine

Admitting

Serving

process starts,
memory empty

read disk snapshot /
ask a warm peer

no peer answered

wait out the quiet period
(Eureka default 5 min)
while clients re-register

state rebuilt
(conservative approximation)

heartbeats / KeepAlives only,
then session ops (Chubby steps 4-8)

Cold

Syncing

Quarantine

Admitting

Serving

Chubby's failover protocol and Eureka's empty-boot behaviour are the same ladder with different rungs: sync what you can, admit heartbeats before work, and only then serve. Sources: Chubby, OSDI '06; Eureka wiki.
Diagram source
03

The decisions that matter

Each of these was argued in public, with a rejected side and a stated reason. The flips-when column is the part to carry into your own review.

Decision: when a component boots empty, may it serve what it has?

Chosen
  • Refuse or restrict until synced. Eureka waits 5 minutes by shipped default when no peer can seed it; Chubby's new master admits KeepAlives before any session operation.
  • Stated reason: partial answers skew all traffic onto whatever happens to have registered first, "causing capacity issues" (Eureka wiki).
Rejected
  • Serve the registry you have, immediately. Faster visible recovery, and honest for data where absence is distinguishable from emptiness.
Flips when
  • Downstream consumers can tell "empty because cold" from "empty because gone". If the API can mark its answer as partial, serve early; if an empty answer reads as truth, silence is the safer lie.

Decision: where does reconnect dispersal live?

Chosen
  • In the client contract. gRPC mandates exponential backoff with jitter and requires that simultaneous backoffs disperse; Consul staggers its own agents' checks since 0.5.0 (2015).
Rejected
  • In the load balancer. Envoy's slow start can ramp traffic into cold endpoints, but its own docs state it is "not very effective" when every endpoint is new, which is the restart case.
Flips when
  • You do not control the clients. Then the balancer and the admission layer are the only levers left, and you should expect to pay for the difference in capacity, as GitLab's 3.3x reconnect-burst autoscale shows.

Decision: on reconnect, does a client resync by re-reading everything?

Chosen
  • Resume from durable progress instead. KEP-1904 initialises the restarted server's cache so watchers can resume; KEP-3157 moves the remaining full resyncs onto the streaming path, taking the worked example from 2.5 TB of transient memory to 2 GB.
Rejected
  • Full LIST on every reconnect. Simple and correct, and the measured cost is O(5x the response) in transient server memory, which is how 16 synchronized informers killed the test apiserver.
Flips when
  • Dataset and client count are small enough that resync-times-five fits in memory. The full re-read is the right simple design below that line; the Kubernetes project only paid for streaming once clusters crossed it.

Decision: may the platform run on itself?

Chosen
  • Keep a statically-booted rung under every self-hosted layer. kubeadm kept static pods for the control plane and never moved etcd off them.
Rejected
  • Full self-hosting (the control plane as DaemonSets scheduled by the control plane). Elegant, upgradeable with the platform's own tools, and unable to survive a full reboot: the pivot deletes the static manifests that would have booted it.
Flips when
  • Never for the bottom layer. Self-hosting is fine for any layer whose restart can be served by a layer below it; it is only the lowest layer that must boot from files and static config alone.

Figure 4 · What a cold component may do, as a decision tree

yes: partial cold

yes

no, but traffic
can be shaped

no: total cold

yes

no

Is anything
still warm?

Warm peer
of the same tier?

Copy from the peer:
mcrouter WarmUpRoute,
Eureka peer sync

Ramp the cold member:
Envoy slow start window

Quorum / durable
state intact?

Boot from disk, gate admission:
heartbeats first, work later

Manual rung: snapshot restore,
new cluster identity, human in the loop

Clients disperse themselves:
mandated backoff + jitter

yes: partial cold

yes

no, but traffic
can be shaped

no: total cold

yes

no

Is anything
still warm?

Warm peer
of the same tier?

Copy from the peer:
mcrouter WarmUpRoute,
Eureka peer sync

Ramp the cold member:
Envoy slow start window

Quorum / durable
state intact?

Boot from disk, gate admission:
heartbeats first, work later

Manual rung: snapshot restore,
new cluster identity, human in the loop

Clients disperse themselves:
mandated backoff + jitter

The tree's first split is the one the warm-up tooling hides: partial cold has engineered paths, total cold has only dispersal, admission and disk. Sources as in the decision blocks above.
Diagram source
DecisionChosenRejectedBecauseEvidence
Serve while cold?Quarantine until syncedServe immediatelyPartial registry skews all traffic onto a few instancesEureka wiki
Dispersal locationClient contract (backoff + jitter)Balancer rampSlow start is inert when every endpoint is coldEnvoy docs, gRPC protocol
Resync shapeResume / streamFull LIST per reconnectO(5x) transient memory; 16 informers killed the serverKEP-3157
Platform on itselfStatic bootstrap rung keptFull self-hosting"Cannot recover from a reboot... without manual intervention"kubeadm docs
Probe the foundation?Metrics and logsReadiness probe on kube-proxyA probe-driven restart of the fixer is itself the herdPR #111857, closed unmerged
Raft log durability bookkeepingbbolt + NoFreelistSync optionLegacy boltdb defaultsFreelist sync cost surfaced at scale after October 2021Consul PR #11720

The fifth row deserves a sentence, because it is the rare rejected pull request whose rejection reasoning is the whole topic in miniature. A contributor proposed a readiness probe for kube-proxy, the component that programs every node's service routing. The reviewer's argument against, quoted in full in the ledger, is that under CPU pressure the health endpoint flaps, the probe then restarts kube-proxy everywhere at once, and "kube-proxy trying to bootstrap from scratch and all the pods on the nodes reconnecting to the Services we just broke" takes the cluster down. The PR closed unmerged in August 2022. Health-checking is not free at the foundation layer: a probe is a restart trigger, and a restart of the foundation is a cold start of everything above it.

04

What broke in production

Four failure classes recur: the synchronized client, recovery amplification, the cold-cache aftershock, and the fixer trapped inside the failure. Each card ends with the rule to carry forward.

Figure 5 · The restart herd, step by step

Backing store (etcd)API server(restarted)Watchers (kubelets,controllers)Backing store (etcd)API server(restarted)Watchers (kubelets,controllers)boots with emptywatch historycheap path gone,fall back to full LIST~5x response size heldin memory per requestevery retry re-synchronizesthesurvivors on the next bootresume watch at revision R410 Gone (R outside historywindow)LIST everything (synchronized)quorum reads, full datasetOOM / crash under the herd
Backing store (etcd)API server(restarted)Watchers (kubelets,controllers)Backing store (etcd)API server(restarted)Watchers (kubelets,controllers)boots with emptywatch historycheap path gone,fall back to full LIST~5x response size heldin memory per requestevery retry re-synchronizesthesurvivors on the next bootresume watch at revision R410 Gone (R outside historywindow)LIST everything (synchronized)quorum reads, full datasetOOM / crash under the herd
The sequence Kubernetes documented in KEP-1904 and KEP-3157: the server's restart erases the history clients need to resume cheaply, so the cheap path collapses into the expensive one at the worst moment. The 5x figure is the project's own measurement, 2022.
Diagram source
Postmortem

Five minutes of resets, an hour of triple fleet (GitLab, 2026)

AssumptionWebSocket capacity is sized by concurrent sessions, which change slowly.
What happenedSomething reset established connections cluster-by-cluster for five minutes (03:00 to 03:05 UTC); the trigger was never found, but every reset client reconnected at once.
Blast radiusAbout 144,000 sessions; error rate 16.14% on the SLI; autoscaling grew the pod fleet from about 100 to 335 to absorb the reconnects.
FixRecovered without mitigation; the durable lesson is the capacity ratio itself.
Design ruleLong-lived connections convert any blip into a herd. Budget roughly 3x steady capacity for reconnect bursts, or meter reconnects at admission.
Postmortem

Deferred work arrives as one batch (GitLab, 2016)

AssumptionPausing a background job (git gc) is free, because it can run later.
What happenedThe counter that triggers gc kept incrementing while gc was disabled during an earlier incident; the next deploy re-enabled the check and every project crossed the threshold at once.
Blast radiusA thundering herd of gc across the fleet immediately after the 8.12 RC7 deploy.
FixReset the accumulated counters; later guards cap how much deferred work re-enables at once.
Design ruleDisabling a consumer does not disable its producer. Anything you pause during an outage is a batch you are scheduling for the recovery.
Design record

Sixteen synchronized informers kill the API server (Kubernetes, 2022)

AssumptionLIST is a read, and reads are cheap enough to serve concurrently.
What happenedEach cache-priming LIST holds about five times the response size in transient memory; sixteen concurrent informers over 400 MB of secrets exhausted the server. Moving roughly 400 GB to prime caches took about 25 minutes.
Blast radiusReproduced loss of the apiserver; the KEP names large-cluster recovery, with every kubelet re-listing pods, secrets and configmaps, as the real-world worst case.
FixKEP-3157 streams the initial state through the watch path: the worked sizing falls from 2.5 TB of transient memory to 2 GB.
Design ruleSize the resync path, not the steady path: peak memory is concurrent resyncers times response size times the serialization multiplier.
SourceKEP-3157
Postmortem

The cold cache as aftershock (GitLab, 2017)

AssumptionOnce yesterday's outage is closed, today starts from normal.
What happenedRecovery from the previous day's incident brought up Redis with an empty cache; the next morning's ordinary traffic, served at cold hit rates, drove 60-second worker timeouts across the web fleet.
Blast radiusSite slow to unresponsive on 2017-09-26; web fleet grown from 7 to 11 nodes as the stopgap.
FixCapacity until the cache warmed; the general fix is a warm route of the mcrouter kind, which needs a warm pool to exist.
Design ruleAn incident is not over when status turns green; it is over when hit rates are back. Track cache warmth as an explicit recovery metric.
Postmortem

The repair tools shared the failed fabric (Google, 2019)

AssumptionDiagnosing fast is the hard part; applying the fix is mechanical.
What happenedA config meant for a few servers hit several regions and removed over half their network capacity. Detection took seconds; the restore took hours because "the same network congestion... also slowed the engineering teams' ability to restore the correct configurations".
Blast radiusMulti-region; YouTube measured a 2.5% drop of views for an hour, Google Cloud Storage a 30% traffic reduction.
FixA post-incident sprint to guard "the entire class of issues", including the tooling path.
Design ruleWalk the path your fix travels while imagining the outage in place. If the fix rides the thing it fixes, you have a cycle, and the outage owns your tools.
Source record

70,000 subscriptions, and the wake-ups cost more than the work (Consul, 2022)

AssumptionFan-out cost scales with the events delivered.
What happenedOn busy clusters, publishing one event scheduled every subscriber's goroutine, and "the majority of scheduled goroutines have no real work to do"; contention concentrated in the kernel's futex layer on high-core-count servers.
Blast radiusBenchmarked at 70,000 concurrent subscriptions on a 128-core server; shipped as a fix in Consul 1.11.3 alongside the post-October-2021 raft storage changes (bbolt, NoFreelistSync).
FixSplit the event buffer by key so only relevant subscribers wake.
Design ruleAt reconnect scale, count wake-ups, not messages. A no-op notification times the whole herd is still a storm.

Figure 6 · The aftershock pattern: recovery schedules the next incident

Day 0Incident declaredMitigation changesstate (cache rebuiltempty, backgroundjob paused)Day 0, laterStatus page greenDeferred work andcold state remain,invisibleDay 1Ordinary trafficmeets cold cache(GitLab 2017-09-26)Deploy re-enablespaused work as onebatch (GitLab #64)Day 1, laterSecond incident,attributed to the firstone's recoveryOutage, recovery, aftershock
Day 0Incident declaredMitigation changesstate (cache rebuiltempty, backgroundjob paused)Day 0, laterStatus page greenDeferred work andcold state remain,invisibleDay 1Ordinary trafficmeets cold cache(GitLab 2017-09-26)Deploy re-enablespaused work as onebatch (GitLab #64)Day 1, laterSecond incident,attributed to the firstone's recoveryOutage, recovery, aftershock
Two GitLab incidents, nine years apart, with the same shape: the recovery action (a cold cache, a paused job) is itself the cause of the next day's incident. Sources: #250, #64.
Diagram source

One class is conspicuously thin in the reachable record: no public account measures a full cold start end-to-end, from power-on to steady-state hit rates, with enough numbers to plan against. The component-level evidence above is the closest thing. Treat your own first full-restart drill as the measurement that does not exist publicly, and instrument it accordingly.

05

Numbers you can plan against

Everything quantitative in the corpus, with context and dates. Measured unless marked otherwise.

MetricValueAtContextAs ofSource
Reconnect backoff constants1s · ×1.6 · cap 120s · jitter ±20%gRPCMandated for every compliant client2017+protocol doc
Transient memory per resync LIST~5× response sizeKubernetesMeasured on the serving path, etcd-backed LIST2022KEP-3157
Synchronized clients to kill one server16 informersKubernetes400 secrets of 1 MB each; synthetic reproduction2022KEP-3157
Cache-priming data moved~400 GB in ~25 minKubernetesSame test; network cost of re-listing2022KEP-3157
Worked resync sizing, before and after streaming2.5 TB → 2 GBKubernetes512 watchers × 400 MB dataset (derived in the KEP: 512×500×2MB×5)2022KEP-3157
Reconnect burst fleet ratio~100 → 335 podsGitLab144,000 sessions reset over 5 minutes2026-10incident #23063
Cold-cache stopgap capacity7 → 11 web nodesGitLabDay-after outage served on an empty Redis2017-09incident #250
Empty-boot quarantine5 minNetflix EurekaRegistry boots with no warm peer; shipped defaultcurrent mastersource
Self-preservation threshold85% renewals / 15 minNetflix EurekaBelow it, the registry stops expiring instanceswiki, undatedwiki
Failover grace period / lease extension45s / 12sGoogle ChubbyClients ride out a master failover without losing sessions2006OSDI '06
Clients per coordinator90,000Google Chubby"clients can overwhelm the master by a huge margin"2006OSDI '06
Subscription fan-out benchmark70,000 subs / 128 coresHashiCorp ConsulEvent-buffer contention fix validation2022-01PR #12080
Slow-start minimum weight10%EnvoyFloor for a cold endpoint's share during the ramp windowcurrent docsdocs
Recovery blast radius2.5% YouTube views (1h); 30% GCS trafficGoogleMulti-region config error; slow restore through congested fabric2019-06blog
Read these carefully

The Kubernetes figures are the project's own synthetic measurements, not a customer incident; the 2.5 TB row is derived arithmetic stated in the KEP, not an observation. The Chubby constants are two decades old and prove the shape, not current values. Nobody in this corpus publishes an end-to-end cold-start duration for a full production stack; the Cloudflare November 2023 account reportedly does, but its host was unreachable from this session and it is deliberately not cited here.

06

The evidence wall

Every source behind this page, graded. Four postmortems, four source records, five design records, three operator docs, two wikis, two papers. No talks: the venues that host them were unreachable from this session's network, and the ledger says so rather than padding the tier.

Postmortem GitLab2026-10

websocketsServices error rate at 16.14% on frontend main

A five-minute burst of connection resets became a reconnect herd: 144,000 sessions, autoscaling from ~100 to 335 pods. The trigger was never identified; the cost ratio was measured anyway.

Carry forwardReconnect capacity is ~3x steady capacity, or you meter at admission.
gitlab.com/gitlab-com/gl-infra/production #23063
Postmortem GitLab2016-09

pushes_since_gc caused a thundering herd

A guard disabled during one incident let a trigger counter accumulate fleet-wide; the next deploy re-enabled the check and scheduled all the deferred work at once.

Carry forwardPausing a consumer schedules a batch for the recovery.
gitlab.com/gitlab-com/gl-infra/production #64
Postmortem GitLab2017-09

Unicorn timeouts on the web fleet

The day after an outage, ordinary traffic on a cold Redis cache drove worker timeouts site-wide; capacity (7 to 11 web nodes) was the stopgap while the cache warmed.

Carry forwardAn incident ends when hit rates recover, not when status turns green.
gitlab.com/gitlab-com/gl-infra/production #250
Postmortem Google2019-06

An update on Sunday's service disruption

Detection took seconds; restoration took hours, because the congestion being fixed also carried the engineers' tooling. YouTube lost 2.5% of views for an hour, GCS 30% of traffic.

Carry forwardIf the fix rides the thing it fixes, the outage owns your tools.
cloud.google.com blog
Decision record Kubernetes2020

KEP-1904: Efficient watch resumption after kube-apiserver reboot

Names the restart, not the crash, as the recurring trigger: a rebooted server has empty watch history, so every returning client is forced into a full relist.

Carry forwardPersist enough progress that clients can resume across your restart.
kubernetes/enhancements KEP-1904
Decision record Kubernetes2022

KEP-3157: watch-list (streaming initial state)

Measures the resync path: ~5x response size of transient memory per LIST, a server lost at 16 concurrent informers, and large-cluster recovery named as "a very expensive storm of LISTs". The fix streams initial state: 2.5 TB to 2 GB in the worked example.

Carry forwardSize the resync path: resyncers × response × serialization multiplier.
kubernetes/enhancements KEP-3157
Decision record Kubernetes2019

KEP-1040: Priority and Fairness for API server requests

The motivation list is a taxonomy of overload at the door, led by "Self-Maintenance crowded out": without explicit priority, the requests that keep the system alive lose to the herd.

Carry forwardReserve capacity for self-maintenance traffic before fairness for the rest.
kubernetes/enhancements KEP-1040
Decision record gRPC (Google)undated

gRPC Connection Backoff Protocol

The client-side half of restart survival as a compliance document: exponential backoff, stated constants, and the requirement that simultaneous backoffs disperse.

Carry forwardPut dispersal in the client contract; the server cannot add it later.
grpc/grpc connection-backoff.md
Decision record Kubernetes2017

kubeadm design v1.9: the self-hosting pivot

The procedure that made the control plane self-hosted ends by deleting the static-pod manifests, which is precisely the bootstrap rung a reboot would need.

Carry forwardRemoving the manual boot path is a one-way door; keep the rung.
kubernetes/kubeadm design_v1.9.md
Source Kubernetes2022-08

PR #111857: readiness probe for kube-proxy (closed, unmerged)

Rejected because a flapping health endpoint would restart the routing layer everywhere at once: "kube-proxy trying to bootstrap from scratch and all the pods on the nodes reconnecting... will bring the whole cluster down".

Carry forwardA probe is a restart trigger; at the foundation, restarts are herds.
kubernetes/kubernetes #111857
Source HashiCorp2015-02

Consul CHANGELOG 0.5.0: random stagger for agent checks

"Random stagger for agent checks to prevent thundering herd": the dispersal rule, shipped eleven years ago, in the changelog of the system at the centre of the best-known total-restart incident.

Carry forwardSynchronized periodic work is a restart herd on a timer.
hashicorp/consul CHANGELOG.md
Source HashiCorp2021-12

Consul PR #11720: bbolt and NoFreelistSync

Weeks after October 2021, Consul's raft log store moved to bbolt and exposed freelist-sync control, with new write-capacity metrics. The reachable artefact of the unreachable postmortem.

Carry forwardThe coordinator's storage bookkeeping is on the recovery path too.
hashicorp/consul #11720
Source HashiCorp2022-01

Consul PR #12080: split the event buffer by key

At 70k subscriptions on 128 cores, publishing one event woke every subscriber; "the majority of scheduled goroutines have no real work to do". Contention lived in the kernel's futex table, not the application locks.

Carry forwardCount wake-ups, not messages, when the herd reconnects.
hashicorp/consul #12080
Vendor docs Kubernetesv1.19 docs

kubeadm self-hosting caveats (archived)

"A self-hosted cluster cannot recover from a reboot of the control-plane node without manual intervention", stated as a product caveat on the feature that removed the static bootstrap rung.

Carry forwardThe bottom layer boots from files, or someone gets paged to be the file.
kubernetes/website (release-1.19)
Wiki Netflixundated

Understanding Eureka peer-to-peer communication

On boot the registry syncs from a neighbour; failing that it quarantines itself, because partial information "skew[s] traffic only to a group of instances and caus[es] capacity issues". The quarantine is a shipped constant, waitTimeInMsWhenSyncEmpty, defaulting to 1000 × 60 × 5.

Carry forwardAn empty answer served as truth is a routing decision you did not mean to make; make the distrust window an explicit, tunable number.
Netflix/eureka wiki
Wiki Metaundated

mcrouter: cold cache warm up setup

The operator's warm-up: misses on the cold pool are filled from a designated warm pool, protecting the backing store. The design presumes the warm pool survived.

Carry forwardWarm-from-peer is the cheap path; have a plan for the day no peer is warm.
facebook/mcrouter wiki
Vendor docs Envoy / CNCFcurrent

Slow start mode

Weight-ramps new endpoints over a window, minimum 10% share, and states its own limit: with every endpoint new, "slow start is not very effective".

Carry forwardBalancer warm-up assumes warm neighbours; total cold needs another lever.
envoyproxy/envoy docs
Vendor docs etcd communityv3.5 docs

Disaster recovery

Below quorum the cluster "disastrously fails"; recovery is a human with a snapshot, and the restored member deliberately loses its identity so it cannot rejoin the old cluster. --force-new-cluster exists and is "strongly discouraged".

Carry forwardThe coordinator's own floor is manual by design; rehearse that runbook.
etcd-io/website recovery.md
Paper Google2006

The Chubby lock service (OSDI '06)

The failover protocol is an admission ladder: rebuild sessions from disk and from the clients themselves, serve KeepAlives before operations. 90,000 clients per master; grace period 45s.

Carry forwardYour clients are a warm source: design the reconnect to carry state back.
chubby-osdi06.pdf (mirror)
Paper AWS2020

Millions of Tiny Databases (NSDI '20)

The recovery coordinator's load profile, stated exactly: near-idle in normal operation, then a latency-critical "burst of work" when a large-scale failure takes many servers offline at once, "most critical at the most challenging time".

Carry forwardRecovery infrastructure is sized by the burst, not the average; its average is misleadingly near zero.
Millions of Tiny Databases (mirror)
07

Build a miniature, then productionise it

Six rungs. The line from toy to real is crossed at rung four, when you start killing everything at once on purpose.

Make the herd visible

One server, 500 client processes holding long-lived connections, no backoff logic. Kill the server for ten seconds, bring it back, graph connections per second and server CPU on reattach.

Done when: you can show the reconnect spike as a multiple of steady connection rate.  Teaches: the restart is a load event, generated by your own clients.

Disperse it in the client

Implement the gRPC backoff algorithm (1s initial, 1.6 multiplier, 120s cap, 20% jitter) in your client. Repeat the kill. Then remove only the jitter and watch the synchronized waves return.

Done when: reattach load spreads across the backoff window, and you have seen why jitter, not backoff, does the spreading.  Teaches: dispersal is the client contract's job.

Add the cold cache

Put a cache in front of a deliberately slow store (add 50ms per read). Restart the cache empty under steady load; measure the amplification at the store and the time back to steady hit rate. Add a warm-peer route of the mcrouter kind and measure again.

Done when: you can state the store's cold-start amplification factor and the warm route's effect on it.  Teaches: the aftershock: the outage is over, the miss storm is not.

Kill everything at once

Now restart server, cache and clients together. The warm route has nothing warm to read from; watch it go inert. Make the server refuse full-resync requests beyond a concurrency cap (admission control) and quarantine itself for N seconds while clients re-register, Eureka-style.

Done when: total cold start completes without the store or the server falling over, slower but monotonically.  Teaches: the warm-peer assumption, by taking the peer away.

Reserve the self-maintenance lane

Give health checks, registrations and leader heartbeats their own priority class at the server's door, with reserved concurrency, per KEP-1040's motivation. Re-run rung four and confirm the system can see itself while overloaded.

Done when: heartbeats hold their latency through the resync storm.  Teaches: fairness is not enough; survival traffic needs a reservation.

Walk the graph, then drill it

Productionising: draw the dependency graph of your real stack with one question per edge, "does this edge exist at boot?". Find the cycle (there is usually one, often through DNS, secrets, or the deploy system). Break it with a static rung, write the black-start runbook, then run a scheduled full-restart drill and measure power-on to steady-state hit rates.

Done when: the drill produces the end-to-end cold-start number nobody publishes.  Teaches: the graph that must be acyclic is the boot-time graph, which is not the runtime graph.

08

Keep hunting

The queries and moves that found this material, in the forms that worked. The wiki-raw and changelog tricks travel to any project.

The restart herd, by its working names

  • kube-apiserver restart LIST storm OOM "thundering herd" relist
  • "reconnect storm" OR "thundering herd" site:gitlab.com gl-infra production
  • "cold start" "black start" postmortem restart dependency cycle bootstrap
  • "waitTimeInMsWhenSyncEmpty"

Design records and rejections

  • repo:kubernetes/kubernetes "thundering herd" is:pr is:closed is:unmerged
  • path:keps "after kube-apiserver reboot" OR "storm of LISTs"
  • grpc "connection backoff" protocol "must ensure" disperse

Changelogs as incident indexes

  • grep -n -i "thundering\|stagger\|freelist\|backpressure" CHANGELOG.md
  • consul changelog "boltdb" freelist bbolt 1.11

Trackers and wikis, fetched raw

  • https://gitlab.com/api/v4/projects/gitlab-com%2Fgl-infra%2Fproduction/issues?search=cold+cache
  • https://raw.githubusercontent.com/wiki/<org>/<repo>/<Page>.md
  • site:github.com wiki "cold cache" warm up mcrouter
09

References

  1. gRPC, Connection Backoff Protocol grpc/grpc repository, undated. Checked 2026-10-07.
  2. Kubernetes, KEP-1904: Efficient watch resumption after kube-apiserver reboot kubernetes/enhancements, 2020. Checked 2026-10-07.
  3. Kubernetes, KEP-3157: Allow informers for getting a stream of data instead of chunking kubernetes/enhancements, 2022. Checked 2026-10-07.
  4. Kubernetes, KEP-1040: Priority and Fairness for API server requests kubernetes/enhancements, 2019. Checked 2026-10-07.
  5. Kubernetes, PR #111857: Enable readiness probe for kube-proxy (closed unmerged) kubernetes/kubernetes, August 2022. Checked 2026-10-07.
  6. Kubernetes, kubeadm: self-hosting the control plane (archived v1.19 docs) kubernetes/website release-1.19 branch. Checked 2026-10-07.
  7. Kubernetes, kubeadm design document v1.9 kubernetes/kubeadm, 2017. Checked 2026-10-07.
  8. HashiCorp, Consul CHANGELOG (0.5.0 stagger entry; 1.11 raft entries) hashicorp/consul. Checked 2026-10-07.
  9. HashiCorp, Consul PR #11720: raft storage to bbolt, NoFreelistSync hashicorp/consul, December 2021. Checked 2026-10-07.
  10. HashiCorp, Consul PR #12080: split event buffer by key hashicorp/consul, January 2022. Checked 2026-10-07.
  11. Netflix, Eureka wiki: Understanding Eureka peer-to-peer communication Netflix/eureka wiki, undated. Checked 2026-10-07.
  12. Netflix, DefaultEurekaServerConfig.java Netflix/eureka, current master. Checked 2026-10-07.
  13. Meta, mcrouter wiki: Cold cache warm up setup facebook/mcrouter wiki, undated. Checked 2026-10-07.
  14. Envoy, Slow start mode (architecture overview) envoyproxy/envoy docs, current. Checked 2026-10-07.
  15. etcd, Disaster recovery (v3.5 operations guide) etcd-io/website, current. Checked 2026-10-07.
  16. Google (Benjamin Treynor Sloss), An update on Sunday's service disruption Google Cloud blog, 2019-06-03. Checked 2026-10-07.
  17. GitLab, production #23063: websocketsServices error rate 16.14% GitLab production tracker, 2026-10-01. Checked 2026-10-07.
  18. GitLab, production #64: pushes_since_gc caused a thundering herd GitLab production tracker, 2016. Checked 2026-10-07.
  19. GitLab, production #250: Unicorn timeouts on web fleet, 2017-09-26 GitLab production tracker, 2017. Checked 2026-10-07.
  20. Mike Burrows, The Chubby lock service for loosely-coupled distributed systems OSDI 2006; PDF mirror fetched and read this session. Checked 2026-10-07.
  21. Brooker, Chen, Ping, Millions of Tiny Databases NSDI 2020; PDF mirror fetched and read this session. Checked 2026-10-07.