Evidence ledger 20 sources Checked 07 Oct 2026

Evidence ledger

One row per claim in Coming back from cold: the restart is a harder load case than the crash: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

Topic: how production systems survive being switched on all at once — why the restart is a harder load case than the failure, where the warm-up machinery the industry built goes inert at total cold, and what has to be true of the dependency graph for a system to boot at all.

Network constraint for this session: outbound access from this container reached four hosts — github.com (HTML pages read via the session's fetch tool; raw file content over raw.githubusercontent.com), gitlab.com (public API and work items), cloud.google.com, and pkg.go.dev. Every artefact cited below was fetched and read in this session from one of those hosts. The canonical postmortems for the best-known total-restart incidents — AWS Kinesis November 2020, AWS DynamoDB/us-east-1 October 2025, Meta October 2021, Roblox October 2021, Cloudflare November 2023 — are hosted on domains this session could not reach and are therefore absent, not overlooked; where the same failure shape is needed, this guide cites the reachable primary record (the Consul changes HashiCorp shipped after October 2021, the Kubernetes project's own design records, GitLab's public incident tracker) and links the sibling guides that cover those incidents from sessions that could reach them. The two papers were read from PDF mirrors hosted in public GitHub repositories and fetched in this session.

All links checked 2026-10-07. One row per claim. Quotes are copied, not paraphrased.

# Org Title Tier Published Checked URL Claim taken from it Supporting quote or figure
1 gRPC (Google) gRPC Connection Backoff Protocol adr undated protocol doc, retrieved 2026-10-07 2026-10-07 https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md The reconnect contract every gRPC client ships: exponential backoff with jitter, and the stated constants "INITIAL_BACKOFF = 1 second; MULTIPLIER = 1.6; MAX_BACKOFF = 120 seconds; JITTER = 0.2"
2 gRPC (Google) gRPC Connection Backoff Protocol adr undated, retrieved 2026-10-07 2026-10-07 https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md Dispersal of simultaneous reconnects is a compliance requirement, not an optimisation "Alternate implementations must ensure that connection backoffs started at the same time disperse, and must not attempt connections substantially more often than the above algorithm."
3 Kubernetes KEP-1904: Efficient watch resumption after kube-apiserver reboot adr 2020, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1904-efficient-watch-resumption/README.md A restarted server boots with empty history, so returning clients cannot resume and must start over "The kube-apiserver watch cache is initialized from etcd at the moment when it starts with empty change history. As a consequence, clients that want to resume a watch immediately after kube-apiserver reboots almost always have a resource version that is out of the history window."
4 Kubernetes KEP-1904 adr 2020, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1904-efficient-watch-resumption/README.md The restart, not the crash, is the recurring trigger of the relist storm "While the inability to reestablish watch doesn't happen often in practice, the kube-apiserver restarts obviously does (e.g. on upgrades). ... all watchers will eventually be forced to relist, causing significant performance and scalability issues for larger clusters."
5 Kubernetes KEP-3157: Allow informers for getting a stream of data instead of chunking adr 2022, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/3157-watch-list/README.md The resync a client performs on reconnect costs the server a five-fold memory amplification per request "The bottom line is around O(5*the_response_from_etcd) of temporary memory consumption. Neither priority and fairness nor Golang garbage collection is able to protect the system from exhausting memory."
6 Kubernetes KEP-3157 adr 2022, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/3157-watch-list/README.md The measured collapse: a server lost to sixteen synchronized cache-priming clients "around 16:40 we lost the server after running 16 informers. During an investigation, we realized that the server allocates a lot of memory for handling LIST requests." (test: 400 secrets of 1 MB each; "it needed ~25 minutes to bring ~400 GB of data across the network!")
7 Kubernetes KEP-3157 adr 2022, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/3157-watch-list/README.md Cluster-wide recovery is the named worst case, and the designed fix moves resync onto the streaming path "in rare cases ... recovery of large clusters with therefore many kubelets and hence informers for pods, secrets, configmap can lead to a very expensive storm of LISTs." Worked sizing: "512 watches of 400mb data: 512*500*2MB*5=2.5TB ↘ 2 GB"
8 Kubernetes KEP-1040: Priority and Fairness for API Server Requests adr 2019, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1040-priority-and-fairness/README.md Under overload, the traffic that keeps the system alive must be explicitly reserved or it is crowded out "Self-Maintenance crowded out. Some requests are for system self-maintenance, such as: node heartbeats, kubelet and kube-proxy work on pods ... In an overload scenario today there is no assurance of priority for these self-maintenance requests."
9 Kubernetes PR #111857: Enable readiness probe for kube-proxy (closed, unmerged) source 2022-08, closed 2022-08-26 2026-10-07 https://github.com/kubernetes/kubernetes/pull/111857 The rejected probe: health-checking a foundational component risks restarting it, and its restart is the herd aojea: "if you start restarting kube-proxy at this point, you can break everything and it will create a thundering herd problem: kube-proxy trying to bootstrap from scratch and all the pods on the nodes reconnecting to the Services we just broke" — PR closed without merging.
10 Kubernetes kubeadm: Configuring your cluster to self-host the control plane (v1.19 docs, archived) vendor retrieved 2026-10-07 (release-1.19 branch) 2026-10-07 https://github.com/kubernetes/website/blob/release-1.19/content/en/docs/setup/production-environment/tools/kubeadm/self-hosting.md The circular dependency stated as a product caveat: a control plane scheduled by itself cannot reboot "Self-hosting in 1.8 and later has some important limitations. In particular, a self-hosted cluster cannot recover from a reboot of the control-plane node without manual intervention."
11 Kubernetes kubeadm design document v1.9: self-hosting pivot adr 2017-12 era, retrieved 2026-10-07 2026-10-07 https://github.com/kubernetes/kubeadm/blob/main/docs/design/design_v1.9.md The pivot deletes the statically-booted copy once the self-hosted copy runs — removing the bootstrap rung "Remove the static Pod manifest file. The kubelet will stop the original static Pod-hosted component that was running."
12 HashiCorp Consul CHANGELOG, 0.5.0 (2015) source 2015-02 2026-10-07 https://github.com/hashicorp/consul/blob/main/CHANGELOG.md Dispersal of synchronized agent work was being retrofitted as early as 2015 "Random stagger for agent checks to prevent thundering herd [GH-546]"
13 HashiCorp Consul PR #11720: replace boltdb with bbolt, expose NoFreelistSync (merged into 1.11) source 2021-12-02/03 2026-10-07 https://github.com/hashicorp/consul/issues/11720 The storage change shipped weeks after the October 2021 Consul-outage class of failure: swap the raft log store and expose the freelist-sync control "In addition to just swapping out bbolt for bolt this also exposes a configuration to allow for disabling boltdb freelist syncing." (Changelog 1.11: "raft: Use bbolt instead of the legacy boltdb implementation"; "raft: Added a configuration to disable boltdb freelist syncing")
14 HashiCorp Consul PR #12080: split event buffer by key (merged 2022-01-28, shipped 1.11.3) source 2022-01 2026-10-07 https://github.com/hashicorp/consul/pull/12080 At high subscription counts the cost is not the work but the wake-ups: contention collapses the server "On a busy cluster with many services and open streams, when an event is published, the majority of scheduled goroutines have no real work to do"; benchmark: "a single server node on an m6i.32xlarge EC2 instance (128 CPU cores) ... 50 worker nodes, each creating 14 gRPC connections ... 100 streams per connection (for a total of 70k concurrent subscriptions)."
15 Netflix Eureka wiki: Understanding Eureka Peer to Peer Communication blog undated wiki, retrieved 2026-10-07 2026-10-07 https://github.com/Netflix/eureka/wiki/Understanding-Eureka-Peer-to-Peer-Communication On boot the registry copies state from a peer; the warm-up path presumes a warm peer exists "When the Eureka server comes up, it tries to get all of the instance registry information from a neighboring node. If there is a problem getting the information from a node, the server tries all of the peers before it gives up."
16 Netflix Eureka wiki: Understanding Eureka Peer to Peer Communication blog undated wiki, retrieved 2026-10-07 2026-10-07 https://github.com/Netflix/eureka/wiki/Understanding-Eureka-Peer-to-Peer-Communication When no warm peer exists, the server prefers silence to a partial answer, and says why "In the case, where the server is not able get the registry information from the neighboring node, it waits for a few minutes (5 mins) so that the clients can register their information. The server tries hard not to provide partial information to the clients there by skewing traffic only to a group of instances and causing capacity issues."
17 Netflix DefaultEurekaServerConfig.java source current master, retrieved 2026-10-07 2026-10-07 https://github.com/Netflix/eureka/blob/master/eureka-core/src/main/java/com/netflix/eureka/DefaultEurekaServerConfig.java The boot-empty silence is a shipped constant: five minutes "waitTimeInMsWhenSyncEmpty", (1000 * 60 * 5)
18 Meta mcrouter wiki: Cold cache warm up setup blog undated wiki, retrieved 2026-10-07 2026-10-07 https://github.com/facebook/mcrouter/wiki/Cold-cache-warm-up-setup The cold-cache problem and the warm-route remedy, in the operator's own words "Every time a new cache box is added to Memcache infrastructure, it has empty cache. This box is called 'cold'. Every request to this box will end up in 'miss', so clients will spend lots of resources to fill the cache. In the worst case this will affect performance of the entire site"; "when we have a cold box and receive a 'miss', try to get a value from 'warm' box (one with filled cache) and set it back to cold box."
19 Envoy Slow start mode (arch overview) vendor undated docs, retrieved 2026-10-07 2026-10-07 https://github.com/envoyproxy/envoy/blob/main/docs/root/intro/arch_overview/upstream/load_balancing/slow_start.rst The load balancer can ramp traffic into a cold endpoint — and its own docs state the mechanism goes inert when everything is cold "Slow start mode is most effective for cases where few new endpoints come up e.g. scale event in Kubernetes. When all the endpoints are relatively new e.g. new deployment in Kubernetes, slow start is not very effective as all endpoints end up getting same amount of requests."
20 etcd Disaster recovery (op-guide, v3.5) vendor undated docs, retrieved 2026-10-07 2026-10-07 https://github.com/etcd-io/website/blob/main/content/en/docs/v3.5/op-guide/recovery.md The coordinator's own worst case is explicitly manual: below quorum there is no automatic rung "If the cluster permanently loses more than (N-1)/2 members then it disastrously fails, irrevocably losing quorum. Once quorum is lost, the cluster cannot reach consensus and therefore cannot continue accepting updates."
21 etcd Disaster recovery (op-guide, v3.5) vendor undated docs, retrieved 2026-10-07 2026-10-07 https://github.com/etcd-io/website/blob/main/content/en/docs/v3.5/op-guide/recovery.md Restore is designed to refuse continuity: a restored member sheds its identity so it cannot rejoin the old cluster "Restoring overwrites some snapshot metadata (specifically, the member ID and cluster ID); the member loses its former identity. ... the restore must start a new logical cluster." Also: "--force-new-cluster to overwrite cluster membership while keeping existing application data ... this is strongly discouraged because it will panic if other members from previous cluster are still alive."
22 Google An update on Sunday's service disruption (2019-06-02 outage) postmortem 2019-06-03 2026-10-07 https://cloud.google.com/blog/topics/inside-google-cloud/an-update-on-sundays-service-disruption The repair tools lived inside the failure: congestion slowed the engineers un-configuring the congestion "Google's engineering teams detected the issue within seconds, but diagnosis and correction took far longer than our target of a few minutes. Once alerted, engineering teams quickly identified the cause of the network congestion, but the same network congestion which was creating service degradation also slowed the engineering teams' ability to restore the correct configurations, prolonging the outage."
23 Google An update on Sunday's service disruption postmortem 2019-06-03 2026-10-07 https://cloud.google.com/blog/topics/inside-google-cloud/an-update-on-sundays-service-disruption Measured blast radius of the slow recovery "Overall, YouTube measured a 2.5% drop of views for one hour, while Google Cloud Storage measured a 30% reduction in traffic."
24 GitLab 2026-10-01: websocketsServices error rate at 16.14% on frontend main (production #23063) postmortem 2026-10-01 2026-10-07 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23063 A five-minute disconnect burst produced a reconnect herd that needed 3.3x the fleet "Established WebSocket connections were reset for a few seconds from 03:00–03:05 UTC ... About 144,000 sessions were affected, and the reconnect storm increased pod load, with autoscaling growing from about 100 to 335 pods."
25 GitLab pushes_since_gc caused a thundering herd (production #64) postmortem 2016-09 (tracker import 2017) 2026-10-07 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/64 Deferred work accumulates while a guard is off, then arrives as one synchronized batch on the next deploy "Yesterday due to infrastructure#464 we disabled git gc ... However, the pushes_since_gc column continued to increase. After 8.12 RC7 was deployed, lots of projects had pushes_since_gc values > 10, resulting in a thundering herd."
26 GitLab Unicorn timeouts on web fleet, 2017-09-26 (production #250) postmortem 2017-09-26 2026-10-07 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/250 The cold cache is an aftershock: yesterday's recovery caused today's outage, and capacity was the stopgap "We believe that the additional stress on the web nodes may be to lack of caching on redis, as we recently brought up a cold cache which was an outcome of the outage the day before"; mitigation: web nodes increased "from 7 to 11".
27 Google (Mike Burrows) The Chubby lock service for loosely-coupled distributed systems (OSDI '06; PDF mirror fetched this session) paper 2006-11 2026-10-07 https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf A failed-over coordinator rebuilds its session state partly from the clients themselves, behind an admission ladder "the new master must reconstruct a conservative approximation of the in-memory state that the previous master had. It does this partly by reading data stored stably on disc ..., partly by obtaining state from clients"; step 4 of the protocol: "The master now lets clients perform KeepAlives, but no other session-related operations."
28 Google (Mike Burrows) Chubby (OSDI '06; PDF mirror) paper 2006-11 2026-10-07 https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf The client population, synchronized, exceeds what the server can absorb by design margin "we have seen 90,000 clients communicating directly with a Chubby master—far more than the number of machines involved. ... the clients can overwhelm the master by a huge margin." (Grace period: "45s by default"; KeepAlive lease extension "default extension is 12s".)
29 AWS (Brooker, Chen, Ping) Millions of Tiny Databases (NSDI '20; PDF mirror fetched this session) paper 2020-02 2026-10-07 https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf The recovery coordinator's load profile is inverted: near-idle normally, latency-critical burst exactly when everything else is broken "In normal operation it handles little traffic, as replication continues to operate with no need to contact the configuration master. However, when large-scale failures (such as power failures or network partitions) happen, a large number of servers can go offline at once, requiring the master to do a burst of work. This work is latency critical ... It is also most critical at the most challenging time: during large-scale failures."

Tier mix: postmortem 5 rows / 4 artefacts, source 5 rows / 5 artefacts, adr 9 rows / 5 artefacts, vendor 4 rows / 3 artefacts, blog 3 rows / 2 artefacts, paper 3 rows / 2 artefacts. 29 rows over 21 distinct artefacts, 4 hosts (github.com, raw.githubusercontent.com, gitlab.com, cloud.google.com).

Absences, stated:

  • No conference talks. USENIX, KubeCon, InfoQ and YouTube were unreachable from this session's network. The talk tier is absent for that reason, not because the topic lacks talks.
  • The five canonical total-restart postmortems are represented indirectly. AWS Kinesis (November 2020), AWS us-east-1/DynamoDB (October 2025), Meta (October 2021), Roblox (October 2021) and Cloudflare PDX01 (November 2023) are all hosted on unreachable domains. Roblox/Consul October 2021 appears here through the fixes HashiCorp shipped (rows 13–14); the AWS October 2025 and Meta 2021 incidents are covered, with citations, by the sibling guide when-name-resolution-fails (2026-09-04), and control-plane-loss-while-serving by serving-without-the-control-plane (2026-10-06), both in this repository.
  • Only 4 distinct hosts, against the 8 the practice targets, for the same network reason. Breadth of organisations (Google, Kubernetes community, HashiCorp, Netflix, Meta, Envoy/CNCF, etcd community, GitLab, AWS) substitutes where it can.
  • No public account measures a full cold start end-to-end (power-on to steady-state) with enough detail to plan against; the closest reachable evidence is component-level (rows 6, 24). That gap is called out in the guide as the reader's residual risk.