Integration & APIs 08 Oct 2026 27 min read

When every client reconnects at once

How production systems hold millions of long-lived client connections, and how they survive the reconnect storm: client backoff contracts, gateway admission control, bounded resume windows, and the full-resync cliff behind them.

Reconstructs the connection layer from GitLab's public incident tracker (144,000 sessions reset, fleet 100 to 335 pods), the matrix.org thundering-herd postmortem, and the contracts Discord, Netflix, gRPC, XMPP, Matrix and Kubernetes keep in public repositories. After reading, an architect can size a reconnect wave against a published 3.3x multiple, write a client dispersal contract with real constants, and place the boundary between cheap resume and the full-resync path that actually causes the outage.

The finding that surprised me

The operators best at surviving reconnect storms cause disconnects on purpose: Netflix's gateway gives every connection a randomized maximum lifetime and closes it politely, so the fleet's reconnects can never synchronize.

What you get out of it

  • Size the connection tier for the return wave, not steady state: GitLab's published multiple is 3.3x pods for a five-minute, 144k-session burst, trigger never found.
  • The reconnect is rarely the load; the resync behind it is. matrix.org's herd was hundreds of clients times several MB of state each, and the backlog outlived the herd by 10-20 minutes.
  • Client dispersal contracts disagree wildly: gRPC mandates jitter (1s, x1.6, cap 120s, jitter 0.2), Socket.IO ships factor 0.5 with a 5s cap, Phoenix's default has no jitter at all; the worst client in your fleet sets the storm's shape.
  • Every resume protocol is a bounded window (Discord seq, Centrifugo offset+epoch capped at 300 messages, Kubernetes watch cache, XEP-0198 counters); design the out-of-window path as a first-class, load-shedded operation.
  • Scheduled, dithered connection death (Zuul: TTL 1800s minus up to 180s dither, 4s close grace) converts the storm into a permanent gentle drizzle, and is the right choice whenever reconnect is a cheap re-register.

Scope

Why this, now. Realtime features keep multiplying standing connections (GitLab's own estimate: ~4,200 connections per 1 RPS of page traffic) while a fresh October 2026 incident shows the return wave still needs 3.3x the fleet and can arrive with no identifiable trigger.

What it does not cover. Voice/video media transport, mobile push via APNs/FCM, request-retry amplification and cold restarts (covered by sibling guides); Slack's, Netflix Pushy's and WhatsApp's canonical accounts are on hosts unreachable from this session and are named as absences rather than cited.

Open the field guide → Self-contained: it loads nothing at read time, follows your system theme, and prints cleanly.