Every source behind this page, graded. This session's network could reach
four hosts (GitHub, GitLab, raw.githubusercontent, cloud.google.com); the ledger in
sources.md names what is absent because of that and what each absence would have added.
Postmortem
GitLab2026-10
websocketsServices error rate at 16.14% (production #23063)
The cleanest published measurement of a reconnect storm: five minutes of resets,
144,000 sessions, autoscaling from ~100 to 335 pods, and a trigger that survived a full
investigation unidentified.
Carry forwardPlan the connection tier against the return
wave (here 3.3×), and assume some storms will arrive without a root cause.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23063
Postmortem
Matrix.org2017-02
Load problems on the Matrix.org homeserver
Several hundred clients resynced one room simultaneously, several MB of JSON each;
ten minutes of overload, 10-20 more to drain the backlog, a repeat two days later, and
monitoring that crashed on the very 500s it should have paged on.
Carry forwardThe reconnect is not the load; the resync is.
Isolate the resync path from the process holding the connections.
https://github.com/matrix-org/matrix.org/blob/main/content/blog/2017/02/2017-02-17-load-problems-on-the-matrix-org-homeserver.md
Postmortem
GitLab2025-04
Increased errors in websockets, git and web (production #19613)
A routine deployment re-created pods and the error ratio rose across three services
while every displaced connection re-arrived. Closed on rollout completion.
Carry forwardA deploy of the connection fleet is a
scheduled mass disconnect; pace it against spare admission capacity.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/19613
Postmortem
GitLab2026-06
Incident review INC-10719: WebSocket closes abnormally (1006)
A Severity 2, 10.5-hour outage in which the only connection-layer signal was close
code 1006, which by definition carries no reason; the real cause lived in a packaging
change far above the socket.
Carry forwardInstrument the application-level handshake
separately from socket establishment; 1006 alone cannot distinguish your failures.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/22244
Vendor
Discord2026 docs
Gateway documentation (discord-api-docs)
The most complete public reconnect contract: jittered first heartbeat "to prevent too
many clients from reconnecting their sessions at the exact same time", resume via
session_id + seq + a dedicated resume_gateway_url, and identify rationing in
max_concurrency buckets per 5 seconds under a daily session-start budget.
Carry forwardAdmission control at session establishment is
a public, documented contract, not an internal emergency lever.
https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx
Blog
Discord / Elixir2020-10
Real time communication at scale with Elixir at Discord
The stateful-gateway data point: 12M concurrent clients, 26M WebSocket events per
second, sessions as processes, fan-out to 200k+ active users in one community. Post
source read from the elixir-lang site repository.
Carry forwardStateful session processes are what make
resume-with-replay and ordered fan-out possible at this scale.
https://github.com/elixir-lang/elixir-lang.github.com/blob/main/src/content/blog/real-time-communication-at-scale-with-elixir-at-discord.md
Blog
Netflixrepo wiki
Zuul wiki: Push Messaging
Netflix's operating manual for a push-connection fleet: two-level registry with
mandatory TTL, the load balancers that mishandle persistent connections, and the
reconnect dither that randomizes each connection's maximum lifetime.
Carry forwardThe registry TTL and the connection lifetime
are the same number; expiry is the consistency mechanism, not hygiene.
https://github.com/Netflix/zuul/wiki/Push-Messaging
Source
Netflixmaster
PushRegistrationHandler.java (Zuul)
The dither, in code: deadline = TTL minus a random slice minus the grace period; at
the deadline the server sends an application-level goaway and force-closes 4 seconds
later if the client has not gone politely.
Carry forwardAsk the client to close first: a client-initiated
close reconnects with warm state instead of discovering the loss by timeout.
https://github.com/Netflix/zuul/blob/master/zuul-core/src/main/java/com/netflix/zuul/netty/server/push/PushRegistrationHandler.java
ADR
gRPCrepo doc
gRPC Connection Backoff Protocol
The reconnect contract as a compliance document: 1s initial, 1.6 multiplier, 120s
cap, 0.2 jitter, and the requirement that simultaneous backoffs disperse.
Carry forwardWrite your client's reconnect behaviour as a
protocol document with constants, not as a library default someone can override.
https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md
ADR
XMPP (XSF)2008-2025
XEP-0198: Stream Management
Stream resumption standardised in 2008: acknowledgement counters both ways, resume
replays the unacknowledged tail, and duplicates are accepted as unavoidable at this
layer.
Carry forwardResume-with-replay is at-least-once by
construction; deduplication belongs to message ids, not to the transport.
https://github.com/xsf/xeps/blob/master/xep-0198.xml
ADR
Matrix.org2021-12
MSC3575: Sliding Sync (proposal)
The design argument for escaping the full-sync cliff: initial sync "can take minutes"
on large accounts, sync time must become independent of account size, and servers may
discard old position tokens if the degrade path is graceful.
Carry forwardBound the resume window explicitly and design
the out-of-window path as a first-class, load-sheddable operation.
https://github.com/matrix-org/matrix-spec-proposals/blob/kegan/sync-v3/proposals/3575-sync.md
Source
Matrix.orgclosed unmerged
PR #3575 on matrix-spec-proposals
The sliding-sync proposal ran 81 commits over several years and closed without
merging; Synapse changelog entries in the thread record the implementation diverging
into a "simplified" variant incompatible with the written spec.
Carry forwardSync redesigns are multi-year migrations; plan
for the proposal and the deployed protocol to drift during the overlap.
https://github.com/matrix-org/matrix-spec-proposals/pull/3575
Source
Matrix.orgpost-2017
matrix-js-sdk commit 4d46251
The client-side half of the 2017 fixes: keepalive-based recovery and a retry delay of
5000 + random(5000) ms, with the commit message "introduce randomness to minimize
possibility of thundering herds".
Carry forwardJitter usually ships after the first storm;
shipping it before costs one line.
https://github.com/matrix-org/matrix-js-sdk/commit/4d46251b15174be443e183b69cbb4f44555f58b2
Source
Socket.IOmain
socket.io-client manager.ts
The defaults a very large share of deployed WebSocket apps actually run: 1s initial
delay, 5s cap, randomizationFactor 0.5, infinite attempts.
Carry forwardA 5-second cap means the whole herd returns
within ~5s of recovery, forever; raise the cap if your fleet is large.
https://github.com/socketio/socket.io/blob/main/packages/socket.io-client/lib/manager.ts
Source
Socket.IO2024-05
Discussion #5030: graceful shutdown with client reconnection
A production operator asks how to drain a node and steer when clients retry. The
thread is marked Closed Unanswered; the maintainer's only guidance covers which close
method to call.
Carry forwardControlled drain is not a given in your stack;
verify it exists before you depend on rolling restarts.
https://github.com/socketio/socket.io/discussions/5030
Source
Phoenixmain
phoenix socket.js
The default socket reconnect schedule starts at 10 ms, caps at 5 s, and contains no
randomization; channel rejoins cap at 10 s. Dispersal is left entirely to the
application.
Carry forwardAudit your framework's shipped schedule; the
frameworks disagree with each other and with the gRPC contract.
https://github.com/phoenixframework/phoenix/blob/main/assets/js/phoenix/socket.js
Vendor
Centrifugal Labs2026 docs
Centrifugo: stream history and recovery
The recovery protocol in full: offset plus epoch, the recovered flag that routes
clients to the cheap or the expensive path, and a 300-publication default cap on any
single recovery.
Carry forwardExpose "recovered: false" to application code;
the app, not the transport, owns the full-reload decision.
https://github.com/centrifugal/centrifugal.dev/blob/main/docs/server/history_and_recovery.md
Blog
Centrifugal Labs2026-03
Scaling AI token streams with Centrifugo
A current application of the same machinery: history kept in Redis rather than
gateway memory, so gateway restarts preserve recovery; the backend stays stateless and
publishes through the broker.
Carry forwardKeep the replay buffer outside the process
that restarts, or every deploy guarantees a resync storm.
https://github.com/centrifugal/centrifugal.dev/blob/main/blog/2026-03-01-scaling-ai-token-streams-with-centrifugo.md
Paper
Google2006
The Chubby lock service (OSDI '06)
The session-survival design that predates every system above: 90,000 clients per
master, a 45s grace period bridging fail-over, 12s leases that stretch under overload,
and a new master that rebuilds state partly from the clients themselves behind a
KeepAlive-only admission ladder.
Carry forwardDuring recovery, accept heartbeats and defer
everything else; the ladder is what turns a stampede into a queue.
https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf
Paper
AWS2020
Millions of Tiny Databases (NSDI '20)
Names the inverted load profile of every reconnection coordinator: near-idle in
normal operation, a latency-critical burst exactly when large-scale failure hits, "most
critical at the most challenging time".
Carry forwardLoad-test the registry and resync paths at
disaster scale, not steady state; steady state tells you nothing about them.
https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf
ADR
Kubernetes2020
KEP-1904: Efficient watch resumption
The resume window's failure mode at control-plane scale: a restarted server boots
with empty history, returning watchers fall out of the window, and the forced relists
cause "significant performance and scalability issues for larger clusters".
Carry forwardA restart empties the resume window at the
exact moment everyone needs it; persist or warm the window across restarts.
https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1904-efficient-watch-resumption/README.md
ADR
GitLabcurrent
doc/development/real_time.md
A working doctrine for adding load to a connection fleet: treat the connection as
ephemeral, estimate ~4,200 standing connections per 1 RPS of page traffic, and roll out
every new connection behind a percentage feature flag.
Carry forwardNew realtime features are capacity events;
estimate and flag them like one.
https://gitlab.com/gitlab-org/gitlab/-/blob/master/doc/development/real_time.md
ADR
GitLab2020-2021
Epic 355: Websockets on Kubernetes
The infrastructure record of standing up the dedicated fleet: WebSocket traffic
proxied "to separate nodes, isolated from the current Web/API nodes", built
observability-first to learn the connection counts before scaling them.
Carry forwardSeparate the connection fleet before the
first big feature needs it, while the connection count is still small enough to
measure calmly.
https://gitlab.com/groups/gitlab-com/gl-infra/-/work_items/355
Vendor
Google Cloud2021-01
Cloud Run gets WebSockets, HTTP/2 and gRPC bidirectional streams
The serverless constraint stated plainly: "WebSockets streams are still subject to
the request timeouts configured on your Cloud Run service". On such platforms scheduled
churn is not a choice you make; it is one the platform makes for you.
Carry forwardIf the platform caps connection lifetime,
resume cheaply or do not build on long-lived connections there.
https://cloud.google.com/blog/products/serverless/cloud-run-gets-websockets-http-2-and-grpc-bidirectional-streams