RECONNECT STORMS  / field guide
Practitioner field guide · 2026-10-08

When every client reconnects at once

How production systems hold millions of long-lived client connections, and what actually happens in the five minutes after those connections drop together. Reconstructed from GitLab's public incident tracker, the matrix.org postmortems, and the connection contracts that Discord, Netflix, Google and the major realtime frameworks keep in public repositories. After reading it you can size a reconnect wave for your own fleet, write the client contract that keeps it survivable, and decide where your resume window ends and your full-resync cliff begins.

27 primary sources 9 production systems 4 published incidents Evidence through October 2026 Read: 22 min
01

The territory

A promise to tell users the moment something changes requires one open conversation per user. Holding the conversations is cheap. Their correlated ending is not.

3.3×
Fleet growth GitLab's autoscaler needed to absorb one five-minute reconnect burst (100 to 335 pods)
144,000
Sessions reset in that burst, in roughly five minutes, with the trigger never found
26M/s
WebSocket events Discord's gateway pushed to 12M concurrent clients
45s
Grace period Chubby grants every client session so a failed-over master survives their mass return

State the problem without naming a protocol: a service has promised its users that the server will speak first. Chat messages, presence dots, CI job status, collaborative cursors, none of these can wait for the next poll, so the service keeps an open, stateful conversation with every active client. The steady state is strange but manageable: enormous numbers of connections, each nearly idle, each costing a socket, some buffers and a heartbeat. GitLab's own capacity note makes the ratio concrete: each 1 request per second of page traffic adds roughly 4,200 standing WebSocket connections (2022-era estimate, revised as features ship). The load profile is inverted relative to everything else you run: almost no work per connection per second, and all of the work concentrated at the moments connections begin.

That concentration is the subject of this guide. Connections do not end independently; they end together, because the things that end them, a deploy, a load balancer config change, a gateway restart, a network blip, are shared. When they end together, every client comes back together, and three costs that were amortised across hours land in the same second: the TLS and protocol handshake, authentication, and the state catch-up that makes the connection useful again. GitLab's October 2026 incident is the cleanest published measurement: connections reset for roughly five minutes, about 144,000 sessions affected, and the websockets fleet autoscaled from about 100 pods to 335 to absorb the return wave, for a disconnection whose trigger was never identified even after ruling out pod crashes, deploys, node events and the cloud load balancer itself.

The finding that reshaped this guide: the operators who are best at surviving reconnect storms do not try to prevent disconnection. They cause it. Netflix's push gateway assigns every connection a randomized maximum lifetime (registry TTL of 1800 seconds minus a random dither of up to 180 seconds, in the shipped defaults), asks the client to close, and force-closes four seconds later if it will not. The storm cannot synchronize because the fleet's connections never share a birthday. Serverless platforms arrive at the same place by constraint rather than choice: Google's own announcement of WebSockets on Cloud Run notes that streams "are still subject to the request timeouts configured on your Cloud Run service". Across the corpus, the mature position is that the mass reconnect is not an anomaly to prevent but a permanent load profile to meter, and the metering happens on three dials: jitter in the client, admission rationing at the gateway, and a bounded resume window behind it.

This guide covers the connection layer and its reconnect dynamics: gateway fleets, client backoff contracts, session resumption, and the fallback resync path. It deliberately does not cover voice and video media transport, mobile push via APNs and FCM (a different delivery contract), generic request retry amplification, or cold-starting stateless services; the sibling guides on retry storms and on cold restarts cover the last two. Slack's two well-known reconnect postmortems and Netflix's Pushy scale figures live on hosts this research session could not reach; they are named in the ledger as absences, and nothing below depends on their contents.

Figure 1 · The storm has two prices, and only one of them is the handshake

cheap, bounded

MBs per client, DB reads

Deploy or rolling restart

Mass disconnect:
every timer starts at the same instant

LB or network reset

Gateway node loss

Reconnect wave
handshake + auth

Resume hit:
replay small tail from buffer

Resume miss:
full state resync

Steady state

Overload feeds back:
timeouts cause more retries

cheap, bounded

MBs per client, DB reads

Deploy or rolling restart

Mass disconnect:
every timer starts at the same instant

LB or network reset

Gateway node loss

Reconnect wave
handshake + auth

Resume hit:
replay small tail from buffer

Resume miss:
full state resync

Steady state

Overload feeds back:
timeouts cause more retries

Every trigger funnels into the same herd, but the herd splits by what each client must do after connecting: resume from a bounded window, or rebuild state from scratch. The resync branch is what overloaded matrix.org for ten minutes in February 2017.
Diagram source
02

How it is actually built

The same five-part shape appears in every published system: a contracted client, a protocol-aware edge, a dedicated gateway fleet, shared session state with a TTL, and a bounded replay buffer in front of the real source of truth.

Figure 2 · Reference architecture across Discord, Netflix, GitLab and Centrifugo

Shared state

Dedicated connection fleet

Edge

resume miss

Client
backoff + jitter contract
resume token (session, seq)

L4 / WebSocket-aware LB
raised idle timeout

Gateway node
auth, heartbeats,
local connection registry,
deploys decoupled from web fleet

Global registry
client to node, TTL'd

Replay buffer
bounded history: offset + epoch

Source of truth
full-resync path, must shed load

Backends publish events

Shared state

Dedicated connection fleet

Edge

resume miss

Client
backoff + jitter contract
resume token (session, seq)

L4 / WebSocket-aware LB
raised idle timeout

Gateway node
auth, heartbeats,
local connection registry,
deploys decoupled from web fleet

Global registry
client to node, TTL'd

Replay buffer
bounded history: offset + epoch

Source of truth
full-resync path, must shed load

Backends publish events

The components every published system shares. The main divergence point is the gateway box: a stateful session process per client at Discord, a stateless registry entry at Netflix.
Diagram source

Start at the edge, because that is where the first published lesson sits. Netflix's Zuul push documentation is blunt about load balancers: persistent connections "throw off many popular load balancers which cut the connection after some period of inactivity", naming classic ELB, older HAProxy and older Nginx, and the stated fixes are a WebSocket-aware L7 balancer or dropping to plain TCP at layer 4, plus a raised idle timeout either way. An idle-timeout mismatch at the edge is a standing storm generator: it converts your quietest connections into a steady background of forced reconnects.

Behind the edge sits a gateway fleet that does three jobs: terminate and authenticate the connection, heartbeat it, and remember who is connected where. Every published system keeps that fleet separate from the request-serving web tier. GitLab's developer documentation says it directly: on GitLab.com, "WebSocket connections are served from dedicated infrastructure, entirely separate from the regular Web fleet", a decision recorded in the infrastructure epic that proposed proxying WebSocket requests "to separate nodes, isolated from the current Web/API nodes". The isolation is not about CPU. It is about deploy cadence: the web tier redeploys many times a day, and every redeploy of whatever process holds the sockets is a self-inflicted mass disconnect (GitLab's own April 2025 incident, below, is exactly that).

Where the architectures genuinely diverge is what the gateway remembers. Discord's gateway holds a stateful session process per connected client; the Elixir case study describes the gateway "responsible for relaying messages and real-time replication", with fan-out to two hundred thousand active users in a single community, and 26 million events per second to clients at 12 million concurrent users (2020 figures). The session process tracks what the client has seen, which is what makes Discord's resume protocol possible. Netflix's Zuul push takes the opposite position for a push-notification workload: the gateway node keeps only "a local, in-memory registry of all the clients connected to it", and a multi-node cluster adds a second-level global datastore mapping client identity to gateway node, with the explicit requirement that the store support "TTL or automatic record expiry". The TTL is not hygiene; it is the consistency mechanism. A registry entry that can quietly outlive its connection would route pushes into the void, so Netflix bounds the lie with expiry and forces the client to re-register by reconnecting.

The final shared component is the replay buffer, and it is the one that decides whether your reconnect wave is cheap. Centrifugo's design states the two primitives most cleanly: every publication in a channel carries an incrementing offset, and the stream carries an epoch so that "a stale offset is never trusted" after the stream itself has been recreated. On resubscribe the server either fills the gap and returns recovered: true, or refuses and the client falls back to a full reload from the application backend. Kubernetes runs the same design at control-plane scale, and KEP-1904 documents what happens when the buffer resets: a restarted API server starts "with empty change history", returning clients fall out of the window, and "all watchers will eventually be forced to relist", the Kubernetes name for the resync cliff. Same shape, third name: XMPP standardised it in 2008 as XEP-0198 stream management, counters on both sides and replay of the unacknowledged tail, with the honest caveat that replay "might result in duplicates; there is no way to prevent such a result in this protocol".

The gateway fleet

Terminates, authenticates, heartbeats. Dedicated so that web-tier deploys do not disconnect everyone; Zuul closes unauthenticated connections after 8 seconds (default) so anonymous sockets cannot squat.

Runs this way at: GitLab, Netflix

The registry and its TTL

Local map on each node, global client-to-node map across the cluster. The TTL bounds how long a dead entry can mislead the router, and sets the ceiling on connection lifetime (Zuul default: 1800 seconds).

Runs this way at: Netflix; Discord holds it as session process state

The replay buffer and the cliff behind it

Bounded history keyed by position (offset, seq, resourceVersion) plus a stream identity (epoch). Outside the window, the client falls onto the full-resync path, which must be load-shed like any other expensive endpoint.

Runs this way at: Centrifugo, Kubernetes, XMPP

Figure 3 · One connection's life, including the death the server schedules for it

open socket

auth ok, registry write

no auth within 8s

heartbeat each interval

dithered deadline reached,
server sends goaway

client closes, or forced after 4s grace

network drop or node loss

jittered exponential delay

retry

resume hit, tail replayed

resume miss, load from source of truth

Connecting

Registered

Closed

Draining

Backoff

Synced

FullResync

open socket

auth ok, registry write

no auth within 8s

heartbeat each interval

dithered deadline reached,
server sends goaway

client closes, or forced after 4s grace

network drop or node loss

jittered exponential delay

retry

resume hit, tail replayed

resume miss, load from source of truth

Connecting

Registered

Closed

Draining

Backoff

Synced

FullResync

The lifecycle every mature gateway implements. Note the draining state: Netflix's gateway asks the client to close and force-closes four seconds later, at a randomized deadline, so the fleet's reconnects never synchronize. From PushRegistrationHandler.java.
Diagram source
03

The decisions that matter

Five forks recur across every published system. Each has a condition that flips the answer, and most of the conditions are about who controls the client.

One disagreement in the record is worth seeing whole before the individual decisions, because it is the purest case of the same question answered three ways. gRPC treats reconnect dispersal as protocol law: constants in a compliance document, and a sentence making dispersal mandatory for any alternate implementation. Discord treats it as defence in depth: the client contract mandates jitter on the first heartbeat, and the server still rations identifies into 5-second buckets because some client somewhere will misbehave. Phoenix, meanwhile, ships a default schedule with no randomization and a 5-second cap, leaving dispersal to the application. None of these teams is careless; they sit at different points on one axis, how much the operator trusts the client population, and that axis is the first decision below.

Decision: does the gateway hold the session, or only a pointer to it?

Chosen
  • Discord: a stateful session process per client, tracking the event sequence the client has seen, enabling resume-with-replay and ordered fan-out at 26M events/s (2020).
  • Netflix Zuul push: a stateless gateway plus a TTL'd global registry mapping client to node; the message source looks the client up and pushes.
Rejected
  • Each rejected the other's model for its own workload. A registry cannot give Discord per-session ordering and catch-up; session processes would give Netflix per-connection state it does not need to deliver a notification.
Flips when
  • The client needs an ordered, resumable stream: hold state in the gateway. The payload is self-contained notifications: hold a pointer, and let the registry TTL bound staleness.

Decision: who enforces reconnect dispersal, the client contract or the server?

Chosen
  • gRPC writes dispersal into the protocol: backoff 1s, multiplier 1.6, cap 120s, jitter 0.2, and "alternate implementations must ensure that connection backoffs started at the same time disperse".
  • Discord enforces it server-side as well: identifies are rationed in max_concurrency buckets per 5 seconds, and a budget of 1,000 session starts per day for standard apps.
Rejected
  • Trusting library defaults. The defaults disagree wildly: Socket.IO ships 1s initial, 5s cap, factor-0.5 jitter; Phoenix's default socket reconnect has no randomization at all and caps at 5 seconds; matrix-js-sdk added 5-10s randomized retry only after the 2017 storms.
Flips when
  • You ship every client binary: a mandated client contract is enough. Any third-party or long-tail clients: the server must ration admission, because the worst client in the fleet sets the storm's shape.

Decision: let connections live, or kill them on a schedule you control?

Chosen
  • Netflix kills its own connections on purpose: each connection's maximum lifetime is the registry TTL (1800s) minus a random dither (up to 180s), the server asks the client to close first, and force-closes after a 4-second grace. The storm is converted into a permanent, gentle drizzle.
Rejected
  • Letting connections live until something external ends them. External enders are correlated (deploys, LB resets, platform request timeouts on Cloud Run and its peers), so lifetimes synchronize and the next ending is a herd.
Flips when
  • Reconnect implies an expensive resync rather than a cheap re-register, as in Matrix-style stateful sync. Then scheduled churn multiplies your most expensive path, and the right spend is a longer-lived connection plus a bigger resume window.
DecisionChosenRejectedBecauseEvidence
Connection fleet placementDedicated nodes, own deploy cadenceServe WebSockets from the web fleetWeb-tier deploys would disconnect every client; isolation also contains saturationGitLab epic 355
Resume window sizeBounded history, graceful degrade to full syncStore every past position token"Servers should not need to store all past since tokens"; unbounded windows move the cost to the server foreverMSC3575
Recovery capCap replayed messages per resume (default 300)Replay arbitrarily large gaps in one frameA resume that replays unbounded history is a resync wearing a resume's clothesCentrifugo docs
Buffer placementHistory outside the gateway process (Redis)In-process historyGateway restarts then keep recovery working, instead of guaranteeing a resync stormCentrifugal blog, 2026
New realtime featuresFeature-flagged, percentage rollout, connection estimate firstShip and watchEach 1 RPS of page traffic is ~4,200 standing connections; capacity must precede the flag flipGitLab real_time.md
Degraded modeFeatures survive without the socketHard dependency on the connection"Treat the connection as ephemeral"; the socket is an accelerator, not a guaranteeGitLab real_time.md

Figure 4 · Choosing your storm controls

yes

no

cheap

resync

Do you ship every
client implementation?

Mandate the contract:
jittered exponential backoff,
gRPC-style constants

Ration admission server-side:
identify buckets per 5s,
Discord-style session budget

Is reconnect a cheap
re-register, or a resync?

Schedule churn:
dithered max lifetime,
goaway then grace close

Buy window:
bounded replay buffer sized
to cover p99 disconnect gap

Load-shed the full-resync path;
it is your real failure surface

yes

no

cheap

resync

Do you ship every
client implementation?

Mandate the contract:
jittered exponential backoff,
gRPC-style constants

Ration admission server-side:
identify buckets per 5s,
Discord-style session budget

Is reconnect a cheap
re-register, or a resync?

Schedule churn:
dithered max lifetime,
goaway then grace close

Buy window:
bounded replay buffer sized
to cover p99 disconnect gap

Load-shed the full-resync path;
it is your real failure surface

Work top to bottom; every terminal is a concrete mechanism from the evidence wall. The left column is only available when you ship the client; the right column is what remains when you do not.
Diagram source
04

What broke in production

Three failure classes account for the published incidents: the lockstep herd, the resync cliff behind it, and the deploy that is a disconnect in disguise.

Postmortem

Five minutes of resets, 3.3× the fleet, no root cause

AssumptionSteady-state capacity plus autoscaling headroom covers the websockets service.
What happenedServer-side disconnect bursts reset established connections cluster by cluster for five minutes; every affected client reconnected on its own schedule, and the schedules agreed.
Blast radiusAbout 144,000 sessions; error rate 16.14% on the SLI; pods autoscaled from ~100 to 335 before settling.
FixNone available: GitLab ruled out pod crashes, deploys, node events, HAProxy restarts and the cloud LB, and "the exact trigger for the resets is still unknown". The system survived on autoscaling alone.
Design ruleSize the connection tier for the return wave, not the steady state: the published multiple is 3.3× for a five-minute burst. You will not always get a root cause; you will always get the wave.
Postmortem

The herd was survivable; the resync behind it was not

AssumptionClients re-syncing room state on an odd event "has been fine in the past".
What happenedOne unusual event made several hundred mobile clients in the same room resync simultaneously, each resync "calculating and syncing several MB of JSON state to each client".
Blast radiusServer overloaded ~10 minutes; a further 10-20 minutes to clear the backlog after the herd dissipated; a repeat two days later; the monitoring crashed on the 500s it was supposed to report.
FixNaive client resync fixed in both mobile SDKs; the server-side resync path moved to worker processes "so that worst case it can't take out the main synapse process"; randomized retry shipped in matrix-js-sdk.
Design ruleMultiply your reconnect count by your resync payload before an incident does it for you. Isolate the resync path so its saturation cannot take the connection path down with it.
Postmortem

The deploy is a mass disconnect wearing a release tag

AssumptionA routine deployment affects the service being deployed.
What happenedA deployment re-created pods; the connections those pods held all ended at the deploy's pace, and errors rose across websockets, git and web while the herd re-arrived.
Blast radiusSeverity 3; elevated error ratios across three services for the duration of the rollout.
FixThe offending deployment completed; the incident closed on recovery. The structural fix is pacing: roll the connection fleet at the speed the survivors can absorb.
Design ruleYour deploy rate is a disconnect rate. Budget it like one: nodes drained per step × connections per node must stay under the fleet's spare admission capacity, or the deploy is indistinguishable from an outage.
Postmortem

Close code 1006 and ten and a half hours of guessing

AssumptionA failing WebSocket will say why it failed.
What happenedDuo Agent Platform workflows died at startup with abnormal closure code 1006, the WebSocket code that by definition carries no reason; the actual cause (a CLI packaging change) was invisible from the connection layer.
Blast radiusSeverity 2; all customers on the default configuration; 10 hours 33 minutes.
FixRolled out a corrected package; the review's lesson is diagnostic, not architectural.
Design rule1006 is the connection layer shrugging. Instrument the application handshake (first message exchanged after open) separately from socket establishment, or every failure above the socket presents as this one opaque code.

The fourth class has no modern public postmortem, and the gap matters: the recovery coordinator that is idle until the exact moment everything needs it. The AWS Physalia paper describes the shape precisely: the configuration master "handles little traffic" in normal operation, but in large-scale failures "a large number of servers can go offline at once, requiring the master to do a burst of work", work that is "most critical at the most challenging time". Chubby's 2006 design answers the same problem from the session side: a failed-over master rebuilds its session table "partly by obtaining state from clients", and during the grace window it "lets clients perform KeepAlives, but no other session-related operations", an admission ladder in which the returning herd is first acknowledged, then slowly re-served. No published incident in this corpus shows that ladder failing, which means the public record cannot tell you whether yours works; only a deliberately drained fleet can.

Figure 5 · The matrix.org failure path, as a sequence

"State store""Homeserver""Hundreds ofclients""State store""Homeserver""Hundreds ofclients"one unusual event in a busyroom~10 min overload, main processsaturatedherd dissipatesresync room state (all clients,same moment)compute + serialize several MBper client500s (monitoring also crashedon these)retries pile onto the backlogcatch-up continues 10-20 minafter the herd
"State store""Homeserver""Hundreds ofclients""State store""Homeserver""Hundreds ofclients"one unusual event in a busyroom~10 min overload, main processsaturatedherd dissipatesresync room state (all clients,same moment)compute + serialize several MBper client500s (monitoring also crashedon these)retries pile onto the backlogcatch-up continues 10-20 minafter the herd
Reconstructed from the 2017 postmortem: the reconnecting clients were never the bottleneck; the several-MB resync each one triggered was, and the backlog outlived the herd by 10-20 minutes.
Diagram source
05

Numbers you can plan against

Everything quantitative this hunt surfaced, dated and sourced. Shipped defaults are a different kind of truth from measured incidents: they tell you what the vendor believes survivable, not what your fleet will do.

MetricValueAtContextAs ofSource
Reconnect-wave capacity multiple3.3×GitLabPods 100 to 335 for a 5-minute disconnect burst, 144k sessions2026incident #23063
Standing connections per 1 RPS of page traffic~4,200GitLabCrude published estimate used for capacity review of new features2022 est.real_time.md
Concurrent clients / event fan-out12M / 26M per sDiscordAcross the gateway fleet, stateful sessions2020case study
Session starts allowed per day1,000DiscordStandard apps; scales to max(2000, guilds/1000 × 5) for large bots2026 docsgateway.mdx
Identify concurrency bucketper 5 sDiscordmax_concurrency identifies allowed per 5-second window, excess gets Invalid Session2026 docsgateway.mdx
Connection max lifetime / dither / close grace1800 / 180 / 4 sNetflix ZuulShipped defaults: registry TTL, reconnect dither window, forced-close grace2026 repoZuul wiki
Client backoff contract1 s · ×1.6 · cap 120 s · jitter 0.2gRPCMandatory for compliant implementations2026 repobackoff doc
Socket.IO client defaults1 s, cap 5 s, jitter 0.5Socket.IOreconnectionDelay / DelayMax / randomizationFactor, infinite attempts2026 repomanager.ts
Phoenix client defaults10 ms first, cap 5 s, no jitterPhoenixStepped schedule [10,50,100,150,200,250,500,1000,2000] then 5000 ms2026 reposocket.js
Session grace period / lease extension45 s / 12 sGoogle ChubbyFail-over absorption for ~90,000 clients per master; lease grows when overloaded2006OSDI '06
Recovery replay cap300 msgsCentrifugoclient.recovery_max_publication_limit default; larger gaps become full reloads2026 docsdocs
Resync overload / backlog drain~10 min / 10-20 minmatrix.orgSeveral hundred clients × several MB state each; recovery lagged the herd2017postmortem
Read these carefully

The GitLab and matrix.org rows are measured incidents; the Discord and Chubby rows are operator-reported scale; every row labelled "repo" or "docs" is a shipped default, which is a vendor's opinion, not a measurement. The GitLab 4,200-connections-per-RPS figure is called "crude" by its own authors and dates from the first rollout. Per-server connection records (WhatsApp's 2M-per-box era, Phoenix's 2M benchmark) are on unreachable hosts and deliberately not quoted here; treat any such number you find as dated the day it was set.

06

The evidence wall

Every source behind this page, graded. This session's network could reach four hosts (GitHub, GitLab, raw.githubusercontent, cloud.google.com); the ledger in sources.md names what is absent because of that and what each absence would have added.

Postmortem GitLab2026-10

websocketsServices error rate at 16.14% (production #23063)

The cleanest published measurement of a reconnect storm: five minutes of resets, 144,000 sessions, autoscaling from ~100 to 335 pods, and a trigger that survived a full investigation unidentified.

Carry forwardPlan the connection tier against the return wave (here 3.3×), and assume some storms will arrive without a root cause.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23063
Postmortem Matrix.org2017-02

Load problems on the Matrix.org homeserver

Several hundred clients resynced one room simultaneously, several MB of JSON each; ten minutes of overload, 10-20 more to drain the backlog, a repeat two days later, and monitoring that crashed on the very 500s it should have paged on.

Carry forwardThe reconnect is not the load; the resync is. Isolate the resync path from the process holding the connections.
https://github.com/matrix-org/matrix.org/blob/main/content/blog/2017/02/2017-02-17-load-problems-on-the-matrix-org-homeserver.md
Postmortem GitLab2025-04

Increased errors in websockets, git and web (production #19613)

A routine deployment re-created pods and the error ratio rose across three services while every displaced connection re-arrived. Closed on rollout completion.

Carry forwardA deploy of the connection fleet is a scheduled mass disconnect; pace it against spare admission capacity.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/19613
Postmortem GitLab2026-06

Incident review INC-10719: WebSocket closes abnormally (1006)

A Severity 2, 10.5-hour outage in which the only connection-layer signal was close code 1006, which by definition carries no reason; the real cause lived in a packaging change far above the socket.

Carry forwardInstrument the application-level handshake separately from socket establishment; 1006 alone cannot distinguish your failures.
https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/22244
Vendor Discord2026 docs

Gateway documentation (discord-api-docs)

The most complete public reconnect contract: jittered first heartbeat "to prevent too many clients from reconnecting their sessions at the exact same time", resume via session_id + seq + a dedicated resume_gateway_url, and identify rationing in max_concurrency buckets per 5 seconds under a daily session-start budget.

Carry forwardAdmission control at session establishment is a public, documented contract, not an internal emergency lever.
https://github.com/discord/discord-api-docs/blob/main/developers/events/gateway.mdx
Blog Discord / Elixir2020-10

Real time communication at scale with Elixir at Discord

The stateful-gateway data point: 12M concurrent clients, 26M WebSocket events per second, sessions as processes, fan-out to 200k+ active users in one community. Post source read from the elixir-lang site repository.

Carry forwardStateful session processes are what make resume-with-replay and ordered fan-out possible at this scale.
https://github.com/elixir-lang/elixir-lang.github.com/blob/main/src/content/blog/real-time-communication-at-scale-with-elixir-at-discord.md
Blog Netflixrepo wiki

Zuul wiki: Push Messaging

Netflix's operating manual for a push-connection fleet: two-level registry with mandatory TTL, the load balancers that mishandle persistent connections, and the reconnect dither that randomizes each connection's maximum lifetime.

Carry forwardThe registry TTL and the connection lifetime are the same number; expiry is the consistency mechanism, not hygiene.
https://github.com/Netflix/zuul/wiki/Push-Messaging
Source Netflixmaster

PushRegistrationHandler.java (Zuul)

The dither, in code: deadline = TTL minus a random slice minus the grace period; at the deadline the server sends an application-level goaway and force-closes 4 seconds later if the client has not gone politely.

Carry forwardAsk the client to close first: a client-initiated close reconnects with warm state instead of discovering the loss by timeout.
https://github.com/Netflix/zuul/blob/master/zuul-core/src/main/java/com/netflix/zuul/netty/server/push/PushRegistrationHandler.java
ADR gRPCrepo doc

gRPC Connection Backoff Protocol

The reconnect contract as a compliance document: 1s initial, 1.6 multiplier, 120s cap, 0.2 jitter, and the requirement that simultaneous backoffs disperse.

Carry forwardWrite your client's reconnect behaviour as a protocol document with constants, not as a library default someone can override.
https://github.com/grpc/grpc/blob/master/doc/connection-backoff.md
ADR XMPP (XSF)2008-2025

XEP-0198: Stream Management

Stream resumption standardised in 2008: acknowledgement counters both ways, resume replays the unacknowledged tail, and duplicates are accepted as unavoidable at this layer.

Carry forwardResume-with-replay is at-least-once by construction; deduplication belongs to message ids, not to the transport.
https://github.com/xsf/xeps/blob/master/xep-0198.xml
ADR Matrix.org2021-12

MSC3575: Sliding Sync (proposal)

The design argument for escaping the full-sync cliff: initial sync "can take minutes" on large accounts, sync time must become independent of account size, and servers may discard old position tokens if the degrade path is graceful.

Carry forwardBound the resume window explicitly and design the out-of-window path as a first-class, load-sheddable operation.
https://github.com/matrix-org/matrix-spec-proposals/blob/kegan/sync-v3/proposals/3575-sync.md
Source Matrix.orgclosed unmerged

PR #3575 on matrix-spec-proposals

The sliding-sync proposal ran 81 commits over several years and closed without merging; Synapse changelog entries in the thread record the implementation diverging into a "simplified" variant incompatible with the written spec.

Carry forwardSync redesigns are multi-year migrations; plan for the proposal and the deployed protocol to drift during the overlap.
https://github.com/matrix-org/matrix-spec-proposals/pull/3575
Source Socket.IO2024-05

Discussion #5030: graceful shutdown with client reconnection

A production operator asks how to drain a node and steer when clients retry. The thread is marked Closed Unanswered; the maintainer's only guidance covers which close method to call.

Carry forwardControlled drain is not a given in your stack; verify it exists before you depend on rolling restarts.
https://github.com/socketio/socket.io/discussions/5030
Source Phoenixmain

phoenix socket.js

The default socket reconnect schedule starts at 10 ms, caps at 5 s, and contains no randomization; channel rejoins cap at 10 s. Dispersal is left entirely to the application.

Carry forwardAudit your framework's shipped schedule; the frameworks disagree with each other and with the gRPC contract.
https://github.com/phoenixframework/phoenix/blob/main/assets/js/phoenix/socket.js
Vendor Centrifugal Labs2026 docs

Centrifugo: stream history and recovery

The recovery protocol in full: offset plus epoch, the recovered flag that routes clients to the cheap or the expensive path, and a 300-publication default cap on any single recovery.

Carry forwardExpose "recovered: false" to application code; the app, not the transport, owns the full-reload decision.
https://github.com/centrifugal/centrifugal.dev/blob/main/docs/server/history_and_recovery.md
Paper Google2006

The Chubby lock service (OSDI '06)

The session-survival design that predates every system above: 90,000 clients per master, a 45s grace period bridging fail-over, 12s leases that stretch under overload, and a new master that rebuilds state partly from the clients themselves behind a KeepAlive-only admission ladder.

Carry forwardDuring recovery, accept heartbeats and defer everything else; the ladder is what turns a stampede into a queue.
https://raw.githubusercontent.com/vcode11/papers/master/chubby-osdi06.pdf
Paper AWS2020

Millions of Tiny Databases (NSDI '20)

Names the inverted load profile of every reconnection coordinator: near-idle in normal operation, a latency-critical burst exactly when large-scale failure hits, "most critical at the most challenging time".

Carry forwardLoad-test the registry and resync paths at disaster scale, not steady state; steady state tells you nothing about them.
https://raw.githubusercontent.com/arpit20adlakha/Computer-Science-Papers-For-System-Design/master/Millions%20of%20Tiny%20Databases.pdf
ADR Kubernetes2020

KEP-1904: Efficient watch resumption

The resume window's failure mode at control-plane scale: a restarted server boots with empty history, returning watchers fall out of the window, and the forced relists cause "significant performance and scalability issues for larger clusters".

Carry forwardA restart empties the resume window at the exact moment everyone needs it; persist or warm the window across restarts.
https://github.com/kubernetes/enhancements/blob/master/keps/sig-api-machinery/1904-efficient-watch-resumption/README.md
ADR GitLabcurrent

doc/development/real_time.md

A working doctrine for adding load to a connection fleet: treat the connection as ephemeral, estimate ~4,200 standing connections per 1 RPS of page traffic, and roll out every new connection behind a percentage feature flag.

Carry forwardNew realtime features are capacity events; estimate and flag them like one.
https://gitlab.com/gitlab-org/gitlab/-/blob/master/doc/development/real_time.md
ADR GitLab2020-2021

Epic 355: Websockets on Kubernetes

The infrastructure record of standing up the dedicated fleet: WebSocket traffic proxied "to separate nodes, isolated from the current Web/API nodes", built observability-first to learn the connection counts before scaling them.

Carry forwardSeparate the connection fleet before the first big feature needs it, while the connection count is still small enough to measure calmly.
https://gitlab.com/groups/gitlab-com/gl-infra/-/work_items/355
Vendor Google Cloud2021-01

Cloud Run gets WebSockets, HTTP/2 and gRPC bidirectional streams

The serverless constraint stated plainly: "WebSockets streams are still subject to the request timeouts configured on your Cloud Run service". On such platforms scheduled churn is not a choice you make; it is one the platform makes for you.

Carry forwardIf the platform caps connection lifetime, resume cheaply or do not build on long-lived connections there.
https://cloud.google.com/blog/products/serverless/cloud-run-gets-websockets-http-2-and-grpc-bidirectional-streams
07

Build a miniature, then productionise it

Six rungs from an evening's toy to a fleet you could defend in a capacity review. The crossing from toy to real is rung 3, where you stop measuring connections and start measuring the storm.

Hold ten thousand connections

A WebSocket echo server plus a load generator opening 10,000 idle connections with heartbeats. Measure per-connection memory and file descriptors at idle.

Done when: you can state bytes per idle connection and your fd ceiling from measurement, not folklore.  Teaches: holding is the cheap half; your numbers will be far below the per-server records for exactly that reason.

Kill the server and watch the lockstep

Restart the server with clients using a fixed reconnect delay. Graph connection attempts per 100 ms. Then add exponential backoff with full jitter and graph again.

Done when: the attempt histogram goes from a spike train to a smear, and you can explain why a fixed delay produced the spikes.  Teaches: the gRPC contract's dispersal clause, from your own graph.

Add a resume protocol

Stamp every message with an offset; keep a bounded ring buffer per channel plus an epoch id. On reconnect, replay the tail when possible and return a recovered flag either way. Kill the server mid-stream and verify no gaps and bounded duplicates.

Done when: clients distinguish recovered:true from false, and a buffer overflow forces the false path visibly.  Teaches: the offset/epoch machinery Centrifugo, Discord and Kubernetes converged on.

Measure the storm properly

Disconnect 100% of a 50k-connection fleet and measure: handshake p99, auth backend load, replay bandwidth, and the capacity multiple versus steady state. Repeat with 10% of clients forced onto the full-resync path.

Done when: you can quote your own 3.3× number and the resync amplification factor.  Teaches: why GitLab's autoscaler needed 235 extra pods, in your own units.

Ration admission and drain politely

Add a token bucket on session establishment (identifies per 5 s), a goaway message asking clients to close, and a dithered maximum connection lifetime. Roll a restart across the fleet and compare error rates with and without the controls.

Done when: a full-fleet rolling restart stays under your SLO with the controls on and breaches it with them off.  Teaches: Discord's max_concurrency and Netflix's dither as cause and effect, not folklore.

Make the fallback path survivable

Put the full-resync endpoint behind explicit load shedding with a client-visible retry-after, keep the replay buffer in an external store, and chaos-test: restart the gateway, the buffer store, and both together.

Done when: a simultaneous gateway+buffer restart degrades to paced full resyncs instead of collapse.  Teaches: the Chubby admission ladder, and why the buffer must outlive the process (KEP-1904's lesson).

08

Keep hunting

The queries that found this material, adapted for a normal network. The incident-tracker and in-repo searches are the highest yield per minute.

Incidents and postmortems

  • site:gitlab.com gl-infra/production websocket reconnect
  • "reconnect storm" OR "thundering herd" postmortem websocket
  • "mass reconnect" incident "status page"
  • slack engineering "websockets" outage 2021

The contracts, in repos

  • repo:grpc/grpc connection-backoff
  • repo:discord/discord-api-docs resume_gateway_url max_concurrency
  • "reconnect" "dither" language:java
  • randomizationFactor OR reconnectAfterMs defaults

Resume protocols and their windows

  • "stream resumption" XEP-0198 implementation notes
  • "watch cache" relist "too old" resourceVersion
  • centrifugo recovery offset epoch "recovered"
  • "sliding sync" MSC3575 "initial sync" minutes

Scale accounts worth reading on a full network

  • "migrating millions of concurrent websockets" envoy slack
  • whatsapp erlang "2 million" connections slides
  • netflix pushy websocket proxy "hundreds of millions"
  • "road to 2 million websocket connections" phoenix
09

References

  1. GitLab, production incident #23063: websocketsServices error rate at 16.14% GitLab.com incident tracker, 2026-10-01. Checked 2026-10-08.
  2. Matthew Hodgson, Load problems on the Matrix.org homeserver matrix.org blog source in GitHub, 2017-02-17. Checked 2026-10-08.
  3. GitLab, Incident Review INC-10719: WebSocket closes abnormally (1006) GitLab.com incident tracker, 2026-06-04. Checked 2026-10-08.
  4. GitLab, production incident #19613: increased errors in websockets, git and web GitLab.com incident tracker, 2025-04-04. Checked 2026-10-08.
  5. Discord, Gateway documentation discord-api-docs repository, current main. Checked 2026-10-08.
  6. José Valim, Real time communication at scale with Elixir at Discord elixir-lang.org blog source in GitHub, 2020-10-08. Checked 2026-10-08.
  7. Netflix, Zuul wiki: Push Messaging Zuul repository wiki. Checked 2026-10-08.
  8. Netflix, PushRegistrationHandler.java Zuul repository, master. Checked 2026-10-08.
  9. gRPC, Connection Backoff Protocol grpc/grpc repository. Checked 2026-10-08.
  10. XMPP Standards Foundation, XEP-0198: Stream Management v1.6.3, 2025-07-28; resumption since 2008. Checked 2026-10-08.
  11. Kegan Dougal, MSC3575: Sliding Sync matrix-spec-proposals, opened 2021-12-20. Checked 2026-10-08.
  12. matrix-spec-proposals PR #3575 (closed, unmerged) GitHub. Checked 2026-10-08.
  13. matrix-js-sdk, commit 4d46251 (randomized reconnection) GitHub. Checked 2026-10-08.
  14. Socket.IO, socket.io-client manager.ts socketio/socket.io monorepo, main. Checked 2026-10-08.
  15. Socket.IO discussion #5030: graceful shutdown with automatic client reconnection GitHub, 2024-05-20, closed unanswered. Checked 2026-10-08.
  16. Phoenix Framework, assets/js/phoenix/socket.js phoenixframework/phoenix, main. Checked 2026-10-08.
  17. Centrifugal Labs, Stream history and recovery centrifugal.dev docs source. Checked 2026-10-08.
  18. Centrifugal Labs, Scaling AI token streams with Centrifugo centrifugal.dev blog source, 2026-03-01. Checked 2026-10-08.
  19. Mike Burrows, The Chubby lock service for loosely-coupled distributed systems OSDI 2006; PDF mirror in a public GitHub repository. Checked 2026-10-08.
  20. Brooker, Chen, Ping, Millions of Tiny Databases NSDI 2020; PDF mirror in a public GitHub repository. Checked 2026-10-08.
  21. Kubernetes, KEP-1904: Efficient watch resumption kubernetes/enhancements, 2020. Checked 2026-10-08.
  22. GitLab, doc/development/real_time.md gitlab-org/gitlab, master. Checked 2026-10-08.
  23. GitLab, Epic 355: Websockets (interactive terminal and Actioncable) on Kubernetes gl-infra epics. Checked 2026-10-08.
  24. Google Cloud, Cloud Run gets WebSockets, HTTP/2 and gRPC bidirectional streams Google Cloud blog, 2021-01-22. Checked 2026-10-08.