advanced 3 min answer

2.4 million live WebSocket sessions sit on a gateway fleet being retired this quarter, with a replacement fleet ready in the same region. Sequence the cutover so no session is forced to reconnect more than once and the subscription path is never asked to absorb a storm.

websocketscutoverdrainingreconnectcapacitysession-resumption
Show the full answer Hide the answer

The sequence

  1. Make the new fleet take new connections only. Both fleets stay behind the same entry point and the new one gets a rising share of new sessions, starting at 1% for a day. Every connection it takes is one the old fleet will never have to drain.
  2. Ship the client behaviour before touching servers. A server-initiated "reconnect after N seconds with jitter" control frame, honoured by a client release, is the difference between a migration and an outage. Measure adoption: if 6% of installed clients are too old to honour it, those sessions are the slow tail and the plan must already account for them.
  3. Verify session resumption works - a resume token plus a server-side cursor, so a reconnecting client replays what it missed rather than refetching state. If this does not exist, build it first; without it every reconnect is a cold start and the arithmetic below gets several times worse.
  4. Drain in waves with a hard rate cap, by shard, watching connect-to-first-message p99 and the auth service's queue depth. Abort the wave on either, rather than on an error rate that appears later.
  5. Let natural churn do most of the work. Mobile sessions turn over in hours. Holding the fleet at new-connections-only for a week retires most of it with zero forced reconnects.

The drain arithmetic

Each reconnect costs a TLS handshake, an authentication call and a subscription restore. If the new fleet sustains roughly 8000 new connections a second and you allocate a third of that to migration, 2400000 sessions divided by 2500 per second is about 16 minutes of continuous drain.

The gateway is rarely the binding constraint. At 20 subscriptions per session, 2500 reconnects a second is 50000 subscribe operations a second against the backplane, and the authentication service sees the same 2500 per second on top of normal traffic. Size the drain against the slowest of those three, not the gateway.

Where it can diverge

A session resumed on the new fleet while the old fleet still believes it holds that session delivers messages twice. Fence with a monotonically increasing session epoch, and have the backplane reject a subscribe carrying a stale epoch. Without the fence, the symptom is duplicate notifications that nobody can reproduce.

The point of no return

Deregistering the old fleet while any shipped client build still resolves it directly, and retiring the old backplane topic namespace. Before those two, everything is reversible by shifting the new-connection share back to zero and letting churn refill the old fleet.

How long it really takes

Realistically a week of patient churn plus a 20-minute forced drain for the long-lived desktop tail, and the honest cost is running two fleets for that week. A plan that promises a 16-minute cutover has quietly assumed every client reconnects on command, which is the assumption that turns a migration into an incident.

Decision rule: drain only what churn will not retire within the window you have, because every forced reconnect is paid for by the authentication and subscription path rather than by the gateway.

Common weak answers

  • "Deploy the new fleet and restart the old one." That is the reconnect storm the question asks you to avoid: 2.4 million sessions returning inside a few seconds, each one authenticating and resubscribing.
  • "The clients will reconnect with backoff." Backoff spreads the retries of a failed reconnect; it does nothing about the first synchronised wave, and in production the first wave is what saturates authentication.
  • "Move DNS to the new fleet." DNS changes where new connections go and does not touch an established WebSocket, so the old fleet still holds every long-lived session when the record's TTL has long expired.