Evidence ledger 24 sources Checked 08 Oct 2026

Evidence ledger

One row per claim in When every client reconnects at once: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

Topic: how production systems hold very large fleets of long-lived client connections (WebSocket, SSE, gRPC streams), and how they survive the moment a large fraction of those connections disconnects and comes back at once: the reconnect storm, the resume protocol, and the full-resync cliff behind it.

Network constraint for this session: outbound access from this container reached four hosts, github.com (HTML pages read via the session's fetch tool; raw file content over raw.githubusercontent.com), gitlab.com (public API and work items), and cloud.google.com. Every artefact cited below was fetched and read in this session from one of those hosts. The best-known reconnect-storm postmortems and accounts hosted elsewhere, Slack's January 4 2021 and May 12 2020 write-ups on slack.engineering, Discord's engineering blog on discord.com, Netflix's "Pushy to the Limit" on netflixtechblog.com, WhatsApp's Erlang Factory 2012 slides on erlang-factory.com, the Zuul Push QCon 2018 talk on archive.qconnewyork.com, Robinhood's March 2020 "thundering herd" letter, and the Phoenix "Road to 2 Million Websocket Connections" benchmark, are hosted on domains this session could not reach and are therefore absent, not overlooked. Where their failure shape is needed, this guide cites the reachable primary record: GitLab's public incident tracker, the matrix.org postmortem whose source lives in a public GitHub repository, and the protocol contracts that Discord, Netflix, gRPC, XMPP and Matrix keep in public repositories. No conference talk was reachable from this session, so the talk tier is empty; that is a constraint of the session, not of the topic. The two papers were read from PDF mirrors hosted in public GitHub repositories and fetched in this session.

All links checked 2026-10-08. One row per claim. Quotes are copied, not paraphrased.

# Org Title Tier Published Checked URL Claim taken from it Supporting quote or figure
1 GitLab Production incident #23063: websocketsServices error rate at 16.14% postmortem 2026-10-01 2026-10-08 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23063 A five-minute disconnect burst produced a reconnect herd needing 3.3x the fleet "Established WebSocket connections were reset for a few seconds from 03:00–03:05 UTC ... About 144,000 sessions were affected, and the reconnect storm increased pod load, with autoscaling growing from about 100 to 335 pods."
2 GitLab Production incident #23063 postmortem 2026-10-01 2026-10-08 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23063 The trigger of a mass disconnect can stay unknown even after a full investigation "Google Cloud found no evidence that the internal passthrough load balancer terminated any connection. We have ruled out pod crashes, deploys, node events, HAProxy restarts, and a sustained Redis issue, but the exact trigger for the resets is still unknown."
3 Matrix.org Foundation Load problems on the Matrix.org homeserver (blog source in public repo) postmortem 2017-02-17 2026-10-08 https://github.com/matrix-org/matrix.org/blob/main/content/blog/2017/02/2017-02-17-load-problems-on-the-matrix-org-homeserver.md The expensive storm is not the reconnect, it is the synchronized state resync behind it "not when you have several hundred clients actively syncing the room, and resulted in a thundering herd effect which overloaded the server for ~10 mins or so whilst they all resynced the room (which, in turn, nowadays, involves calculating and syncing several MB of JSON state to each client)."
4 Matrix.org Foundation Load problems on the Matrix.org homeserver postmortem 2017-02-17 2026-10-08 https://github.com/matrix-org/matrix.org/blob/main/content/blog/2017/02/2017-02-17-load-problems-on-the-matrix-org-homeserver.md Recovery lags the herd: the backlog outlives the burst "The traffic load was then high enough that it took the server a further 10-20 minutes for the server to fully catch up and recover after the herd had dissipated. We then had a repeat performance on Monday morning of the same failure mode."
5 Matrix.org Foundation Load problems on the Matrix.org homeserver postmortem 2017-02-17 2026-10-08 https://github.com/matrix-org/matrix.org/blob/main/content/blog/2017/02/2017-02-17-load-problems-on-the-matrix-org-homeserver.md The structural fix isolates the resync path from the main process "A temporary mitigation is in now place which moves the server-side code to worker processes so that worst case it can't take out the main synapse process and can scale better."
6 GitLab Incident Review INC-10719: Duo workflows fail, WebSocket closes abnormally (1006) postmortem 2026-06-04 2026-10-08 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/22244 A WebSocket-dependent feature was down 10.5 hours; the close code said almost nothing "Duo Agent Platform workflows were failing to start due to abnormal WebSocket connection closures (code 1006)" (Severity 2; "Total Duration: 10 hours, 33 minutes").
7 GitLab Production incident #19613: increased errors in websockets, git and web postmortem 2025-04-04 2026-10-08 https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/19613 An ordinary deploy is a mass-disconnect generator for the connection fleet "This increase in errors was due to a recent deployment that caused pods to be re-deployed."
8 Discord Gateway documentation (developer docs repo) vendor retrieved from main branch 2026-10-08 https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx The first heartbeat is jittered, and the stated purpose is storm prevention "jitter is an offset value between 0 and heartbeat_interval that is meant to prevent too many clients (both desktop and apps) from reconnecting their sessions at the exact same time (which could cause an influx of traffic)."
9 Discord Gateway documentation vendor retrieved from main branch 2026-10-08 https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx Session establishment is rationed server-side in 5-second buckets "max_concurrency | integer | Number of identify requests allowed per 5 seconds"; "Apps are limited by maximum concurrency (max_concurrency in the session start limit object) when identifying. If your app exceeds this limit, Discord will respond with a Invalid Session (opcode 9) event."
10 Discord Gateway documentation vendor retrieved from main branch 2026-10-08 https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx Resume needs exactly three tokens, and resuming connections are steered to a different URL than fresh ones "it will need three values: the session_id and the resume_gateway_url from the Ready event, and the sequence number (s) from the last Dispatch (opcode 0) event it received before the disconnect."; "If your app doesn't use the resume_gateway_url when reconnecting, it will experience disconnects at a higher rate than normal."
11 Discord Gateway documentation vendor retrieved from main branch 2026-10-08 https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx Daily session starts are budgeted; the budget scales with fleet size "The session start limit for these bots will also be increased from 1000 to max(2000, (guild_count / 1000) * 5) per day."
12 Discord / Elixir Real time communication at scale with Elixir at Discord (post source in public repo; published on elixir-lang.org) blog 2020-10-08 2026-10-08 https://github.com/elixir-lang/elixir-lang.github.com/blob/main/src/content/blog/real-time-communication-at-scale-with-elixir-at-discord.md The measured scale of a stateful-session gateway fleet "They have crossed more than 12 million concurrent users across all servers, with more than 26 million WebSocket events to clients per second, and Elixir is powering all of this."
13 Discord / Elixir Real time communication at scale with Elixir at Discord blog 2020-10-08 2026-10-08 https://github.com/elixir-lang/elixir-lang.github.com/blob/main/src/content/blog/real-time-communication-at-scale-with-elixir-at-discord.md The gateway holds sessions as stateful processes, not entries in a lookup table "Elixir was initially picked to power the WebSocket gateway, responsible for relaying messages and real-time replication"; "it is not unlikely to encounter more than two hundred thousand active users in those servers. If someone changes their username, Discord has to broadcast this change to all connected users."
14 Netflix Zuul wiki: Push Messaging blog retrieved from repo wiki 2026-10-08 https://github.com/Netflix/zuul/wiki/Push-Messaging The gateway is split from the registry: local in-memory registry per node, global TTL'd store across nodes "Each Zuul Push server maintains a local, in-memory registry of all the clients connected to it ... In case of a multi-node push cluster, a second level, off-the box global datastore is needed ... the chosen datastore should support ... TTL or automatic record expiry of some sort."
15 Netflix Zuul wiki: Push Messaging blog retrieved from repo wiki 2026-10-08 https://github.com/Netflix/zuul/wiki/Push-Messaging The server deliberately randomizes each connection's maximum lifetime "zuul.push.reconnect.dither.seconds | Randomization window for each client's max connection lifetime. Helps in spreading subsequent client reconnects across time | 180 seconds" (registry TTL default 1800 seconds; client close grace period 4 seconds).
16 Netflix Zuul wiki: Push Messaging blog retrieved from repo wiki 2026-10-08 https://github.com/Netflix/zuul/wiki/Push-Messaging Standard load balancers mishandle persistent connections; the fix is L4 mode or a WebSocket-aware LB "This throws off many popular load balancers which cut the connection after some period of inactivity. Amazon Elastic Load Balancers (ELB) and older versions of HAProxy and Nginx all have this issue."; "increase the IDLE timeout value of your load balancer".
17 Netflix PushRegistrationHandler.java (Zuul source) source current master, retrieved 2026-10-08 https://github.com/Netflix/zuul/blob/master/zuul-core/src/main/java/com/netflix/zuul/netty/server/push/PushRegistrationHandler.java The dither is implemented as a scheduled, randomized, client-first close "private int ditheredReconnectDeadline() { int dither = ThreadLocalRandom.current().nextInt(RECONNECT_DITHER.get()); return PUSH_REGISTRY_TTL.get() - dither - CLIENT_CLOSE_GRACE_PERIOD.get(); }"; "// Application level protocol for asking client to close connection" then "// Force close connection if client doesn't close in reasonable time after we made request".
18 gRPC (Google) gRPC Connection Backoff Protocol adr undated protocol doc, retrieved 2026-10-08 https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md The reconnect contract every gRPC client must ship, with its constants "INITIAL_BACKOFF = 1 second; MULTIPLIER = 1.6; MAX_BACKOFF = 120 seconds; JITTER = 0.2".
19 gRPC (Google) gRPC Connection Backoff Protocol adr undated protocol doc, retrieved 2026-10-08 https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md Dispersal is a compliance requirement, not an optimisation "Alternate implementations must ensure that connection backoffs started at the same time disperse, and must not attempt connections substantially more often than the above algorithm."
20 XMPP Standards Foundation XEP-0198: Stream Management adr v1.6.3, 2025-07-28 (spec source in public repo) 2026-10-08 https://github.com/xsf/xeps/blob/master/xep-0198.xml Stream resumption is a 2008-era standardised design: counters on both sides, replay of the unacknowledged tail "This specification defines an XMPP protocol extension for active management of an XML stream between two XMPP entities, including features for stanza acknowledgements and stream resumption." (resumption added in revision 4, 2008-09-08).
21 XMPP Standards Foundation XEP-0198: Stream Management adr v1.6.3, 2025-07-28 2026-10-08 https://github.com/xsf/xeps/blob/master/xep-0198.xml Resume-with-replay accepts duplicates as a protocol-level cost "Because unacknowledged stanzas might have been received by the other party, resending them might result in duplicates; there is no way to prevent such a result in this protocol, although use of the XMPP 'id' attribute on all stanzas can at least assist the intended recipients in weeding out duplicate stanzas."
22 Matrix.org Foundation MSC3575: Sliding Sync (proposal text) adr opened 2021-12-20, retrieved from proposal branch 2026-10-08 https://github.com/matrix-org/matrix-spec-proposals/blob/kegan/sync-v3/proposals/3575-sync.md The full-resync path grows with account size until login takes minutes "On large accounts with thousands of rooms, the initial sync operation can take minutes to perform. This significantly delays the initial login to Matrix clients, and also makes incremental sync very heavy when resuming after any significant pause in usage."
23 Matrix.org Foundation MSC3575: Sliding Sync (proposal text) adr opened 2021-12-20 2026-10-08 https://github.com/matrix-org/matrix-spec-proposals/blob/kegan/sync-v3/proposals/3575-sync.md Resume windows are deliberately bounded; falling out of the window must degrade gracefully "Servers should not need to store all past since tokens. If a since token has been discarded we should gracefully degrade to initial sync."
24 Matrix.org Foundation PR #3575 on matrix-spec-proposals (closed, unmerged) source 2021-12-20, closed 2026-10-08 https://github.com/matrix-org/matrix-spec-proposals/pull/3575 The sync redesign ran 81 commits and was closed without merging; the argument is recorded in the thread PR state: Closed, not merged; branch kegan/sync-v3; 81 commits. Synapse changelog entries in the thread show the implementation diverged ("the simplified API is slightly incompatible with what's in the current MSC").
25 Matrix.org Foundation matrix-js-sdk commit 4d46251: keepalive-based reconnection source commit by dbkr, retrieved 2026-10-08 https://github.com/matrix-org/matrix-js-sdk/commit/4d46251b15174be443e183b69cbb4f44555f58b2 After the 2017 incidents the client library shipped randomized retry Commit message: "Just use the keepalive logic to recover from lost internet connections."; "Also introduce randomness to minimize possibility of thundering herds." Retry delay in code: 5000 + Math.floor(Math.random() * 5000) ms.
26 Socket.IO socket.io-client manager.ts (defaults) source current main, retrieved 2026-10-08 https://github.com/socketio/socket.io/blob/main/packages/socket.io-client/lib/manager.ts The shipped client defaults: 1s initial delay, 5s cap, factor 0.5 jitter, infinite attempts "reconnectionAttempts @default Infinity; reconnectionDelay @default 1000; reconnectionDelayMax @default 5000; randomizationFactor @default 0.5".
27 Socket.IO Discussion #5030: graceful shutdown with automatic client reconnection source 2024-05-20, closed unanswered 2026-10-08 https://github.com/socketio/socket.io/discussions/5030 Controlled drain with steered reconnect has no supported answer in one of the most-deployed WebSocket stacks Asker: "there is no way to specify whenever the clients that receive the disconnection event should retry to reconnect"; a second user hit the same on AWS autoscaling; thread state: "Closed Unanswered". Maintainer addressed only server.close() vs io.close().
28 Phoenix Framework phoenix socket.js (client defaults) source current main, retrieved 2026-10-08 https://github.com/phoenixframework/phoenix/blob/main/assets/js/phoenix/socket.js A major framework's default reconnect schedule has no randomization and caps at 5 seconds "return [10, 50, 100, 150, 200, 250, 500, 1000, 2000][tries - 1] \|\| 5000" (socket reconnect); channel rejoin: "return [1000, 2000, 5000][tries - 1] \|\| 10000".
29 Centrifugal Labs Centrifugo docs: Stream history and recovery (docs source in public repo) vendor current docs, retrieved 2026-10-08 https://github.com/centrifugal/centrifugal.dev/blob/main/docs/server/history_and_recovery.md Recovery is position (offset) plus stream identity (epoch); a changed epoch invalidates the position "offset — an incremental uint64 stamped on each publication"; "A changed epoch is Centrifugo's way of saying 'this is not the stream you were reading', so a stale offset is never trusted."
30 Centrifugal Labs Centrifugo docs: Stream history and recovery vendor current docs, retrieved 2026-10-08 https://github.com/centrifugal/centrifugal.dev/blob/main/docs/server/history_and_recovery.md The recovered flag routes the client to the cheap path or the full reload, and recovery is capped "If it can, the missed publications come back in the subscribe reply ... and the client sees recovered: true. If it can't, it returns recovered: false and no publications."; "client.recovery_max_publication_limit | Cap on publications recovered in one go (default 300)".
31 Centrifugal Labs Blog: Scaling AI token streams with Centrifugo blog 2026-03-01 2026-10-08 https://github.com/centrifugal/centrifugal.dev/blob/main/blog/2026-03-01-scaling-ai-token-streams-with-centrifugo.md The history buffer lives outside the gateway process so gateway restarts keep recovery working "channel history is stored in Redis, not in Centrifugo's process memory. If you restart Centrifugo while a stream is active, recovery still works because the history survives in Redis."
32 Google (Mike Burrows) The Chubby lock service for loosely-coupled distributed systems (OSDI '06; PDF mirror fetched this session) paper 2006-11 2026-10-08 https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf Session-holding servers see clients far outnumber machines, and leases absorb the mass-return "we have seen 90,000 clients communicating directly with a Chubby master"; "The client waits a further interval called the grace period, 45s by default."; "The default extension is 12s, but an overloaded master may use higher values to reduce the number of KeepAlive calls it must process."
33 Google (Mike Burrows) The Chubby lock service (OSDI '06; PDF mirror) paper 2006-11 2026-10-08 https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf A recovering session server rebuilds state from the clients themselves, behind an admission ladder "the new master must reconstruct a conservative approximation of the in-memory state that the previous master had. It does this partly by reading data stored stably on disc ... partly by obtaining state from clients"; "The master now lets clients perform KeepAlives, but no other session-related operations."
34 AWS (Brooker, Chen, Ping) Millions of Tiny Databases (NSDI '20; PDF mirror fetched this session) paper 2020-02 2026-10-08 https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf The reconnection coordinator's load profile is inverted: idle in steady state, burst-critical in disaster "In normal operation it handles little traffic ... when large-scale failures (such as power failures or network partitions) happen, a large number of servers can go offline at once, requiring the master to do a burst of work."; "It is also most critical at the most challenging time: during large-scale failures."
35 Kubernetes KEP-1904: Efficient watch resumption adr 2020, retrieved 2026-10-08 https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1904-efficient-watch-resumption/README.md A restarted server boots with an empty resume window, so returning clients fall onto the expensive path "The kube-apiserver watch cache is initialized from etcd at the moment when it starts with empty change history. As a consequence, clients that want to resume a watch immediately after kube-apiserver reboots almost always have a resource version that is out of the history window."; "all watchers will eventually be forced to relist, causing significant performance and scalability issues for larger clusters."
36 GitLab doc/development/real_time.md (developer design doc) adr current master, retrieved 2026-10-08 https://gitlab.com/gitlab-org/gitlab/-/blob/master/doc/development/real_time.md Each 1 RPS of page traffic creates roughly 4,200 standing connections, and new connections roll out behind flags "it is possible to crudely estimate that each 1 request per second to a page adds approximately 4200 WebSocket connections."; "ensure that the code establishing the new WebSocket connection is feature flagged and defaulted to off. A careful, percentage-based roll-out".
37 GitLab doc/development/real_time.md adr current master, retrieved 2026-10-08 https://gitlab.com/gitlab-org/gitlab/-/blob/master/doc/development/real_time.md The connection is treated as ephemeral by doctrine, not as an edge case "Treat the connection as ephemeral and ensure the feature you're building is backwards compatible. Ensure critical functionality degrades gracefully when a WebSocket connection isn't available."
38 GitLab Epic 355: Websockets (interactive terminal and Actioncable) on Kubernetes adr retrieved 2026-10-08 https://gitlab.com/groups/gitlab-com/gl-infra/-/work_items/355 The connection fleet is deployed as dedicated infrastructure, separate from the web fleet "it is advisable to proxy WebSocket requests to separate nodes, isolated from the current Web/API nodes"; real_time.md confirms: "WebSocket connections are served from dedicated infrastructure, entirely separate from the regular Web fleet and deployed with Kubernetes."
39 Google Cloud Cloud Run gets WebSockets, HTTP/2 and gRPC bidirectional streams vendor 2021-01-22 2026-10-08 https://cloud.google.com/blog/products/serverless/cloud-run-gets-websockets-http-2-and-grpc-bidirectional-streams On serverless platforms the platform itself caps connection lifetime, making scheduled churn unavoidable "It's worth noting that WebSockets streams are still subject to the request timeouts configured on your Cloud Run service. If you plan to use WebSockets, make sure to set your request timeout accordingly."

Tier mix: postmortem 7 rows / 4 artefacts, source 6 rows / 6 artefacts, adr 10 rows / 7 artefacts, vendor 7 rows / 4 artefacts, blog 6 rows / 4 artefacts, paper 3 rows / 2 artefacts. 39 rows over 27 distinct artefacts, 4 hosts (github.com, raw.githubusercontent.com, gitlab.com, cloud.google.com).

Absences, stated:

  • Slack. The two canonical reconnect-storm postmortems (May 12 2020, January 4 2021) and the Envoy WebSocket migration account live on slack.engineering, unreachable from this session. They are named in the guide as part of the public record but no claim in the guide depends on their contents.
  • WhatsApp and Phoenix benchmark figures. The 2012 Erlang Factory slides (2M+ connections per server) and the 2015 Phoenix 2M-connection benchmark are on unreachable hosts; per-server connection-count records are therefore reported as absent rather than quoted.
  • Netflix Pushy scale figures. "Hundreds of millions of concurrent connections" is on netflixtechblog.com, unreachable; the guide cites only what the Zuul repository itself documents (mechanism and defaults, not fleet size).
  • Talks. No conference talk hosting site was reachable; the talk tier is empty for that reason alone. The Zuul Push QCon 2018 talk and the WhatsApp Erlang Factory 2012 talk are the two a reader should seek out first.