metric

Peak Concurrency

also called Concurrent Peak Load, Simultaneous Users

The maximum number of simultaneously connected or active users a system must hold at once, which drives connection memory and fan-out cost in a way that request-rate planning cannot capture.

capacity planninglive eventsconcurrencyfan-outhotstarwebsockets

A platform sized confidently on 400,000 requests per second falls over at a live sports final. The request rate was correct. The thing it omitted was that 24 million people were holding a connection at the same moment, and every scoreboard update had to reach all of them.

Request rate is a flow; concurrency is a stock. They size different things, and a capacity plan built only on the flow misses two costs entirely: the memory consumed by connections that are merely open, and the fan-out multiplier that turns one event into one message per connected client.

Disney reported 59 million peak concurrent viewers for the 2023 cricket World Cup final, having passed 53 million at the semifinal — the scale at which this distinction stops being academic.

Why it matters

Two costs scale with concurrency rather than with throughput, and the second one dominates:

  • Connection memory. An idle TLS WebSocket costs roughly 30–60 KB of kernel and user-space buffers. At 45 KB, 24 million connections is about 1.1 TB before any application state — perhaps 50 nodes at 500,000 connections each. Affordable, and it means the connection tier must be scaled on a different axis from anything doing work.
  • Fan-out. One broadcast to 24 million connections is 24 million messages. At one update per second that is 24M messages/s, against perhaps 590,000 events/s inbound even during a post-wicket spike. Three orders of magnitude, and it decides the whole architecture.

The third property is the arrival shape. Concurrency at a scheduled event is a step function, not a ramp. Five million users arriving within a minute of kick-off is not a traffic level that autoscaling reacting in 60–90 seconds can meet.

Implementation patterns

  • Compute the fan-out term first. It is the largest number in the plan and the one most under your control. Coalescing updates from once per second to once per five seconds divides the dominant cost by five and no user can perceive the difference.
  • Separate the connection tier. Stateless relays that hold sockets and forward messages, with no business logic and no shared state, scaled on connection count alone.
  • Build the broadcast payload once. At this fan-out there is no budget for per-user computation: the payload is identical for everyone and personalisation happens on the client or not at all.
  • Pre-provision to the predicted peak before the event. This is possible precisely because the schedule is known months ahead, which makes live events easier than organic growth, not harder.
  • A tested degradation ladder. Shed in a defined order — reactions, then poll results, then comment fan-out, then update frequency — leaving the video and the score. Decide it in advance and rehearse it.
  • Plan reconnect storms. A network blip disconnects millions at once and they all reconnect immediately. The reconnect peak can exceed the steady-state peak, so client backoff with jitter is a capacity control rather than a politeness.

Industry example

Disney+ Hotstar's reported concurrency records — 59 million at the 2023 final, 53 million at the India–New Zealand semifinal, with the record broken several times during that tournament — are the public benchmark for this class of problem. The instructive part for an architect is what those numbers imply rather than the numbers themselves: at that concurrency the video is a CDN capacity purchase, and the engineering difficulty sits entirely in the stateful features beside the stream. Concurrency is the metric that separates the easy half of a live platform from the hard half.

Failure scenarios

  • Connection-tier memory exhaustion, where nodes OOM at a connection count well below the theoretical limit because application state was co-located with the sockets.
  • Fan-out saturating the internal message bus, so the broadcast path backs up and every client sees stale state while the system reports healthy throughput.
  • Reconnect storms after a brief network event, where the recovery load exceeds the original peak and the platform cannot get back up.
  • Autoscaling arriving after the moment. The scale-out completes two minutes after the goal, by which time the spike is over and the damage is done.
  • Sizing from average concurrency. A match averaging 20 million and peaking at 59 million is a 3x error, and the peak is the only number that matters.

Trade-offs

Choose Gains Pays
Pre-provision to predicted peak Survives a step-function arrival Paying for capacity idle most of the time
Autoscale reactively Cost tracks usage Useless against arrivals faster than the scaling loop
Aggressive coalescing Divides the dominant fan-out term directly Slightly staler updates, imperceptible in most products

When not to use it

Below roughly a million messages per second of fan-out, none of this architecture is justified. At 100,000 concurrent viewers with 40% connected, 40,000 connections is about 1.8 GB of buffers — a couple of instances — and 40,000 messages/s is comfortable for one managed pub/sub topic behind an off-the-shelf WebSocket gateway. Building the bespoke version earlier buys an operational burden and a distributed system to debug in exchange for headroom you will not use.

Concurrency is also the wrong primary metric for a request/response workload. For a conventional API, requests per second and p99 latency size the system correctly, and concurrency matters only as it bounds the connection and thread pools. Reach for concurrency as the planning metric when connections are long-lived and when one event must reach many of them — streaming, chat, collaborative editing, live dashboards, multiplayer.

The decision rule: compute the fan-out term. If coalescing can hold it under about a million messages per second, buy a managed service and spend the engineering elsewhere.

Interview question

Q: You are asked to support 50 million concurrent viewers with live reactions and a scoreboard. Where do you start, and what do you refuse?

What a strong answer covers: starting with the fan-out arithmetic rather than the connection count, and identifying it as three orders of magnitude above inbound · the connection-memory estimate with its planning number, and separating the connection tier for that reason · refusing per-client computation on the broadcast path, and saying where personalisation moves to · coalescing as the only lever with the right exponent, with the specific trade it makes · pre-provisioning rather than autoscaling, justified by the step-function arrival and the known schedule · a rehearsed degradation ladder · reconnect storms as a distinct peak · and naming the scale below which none of it is warranted.

Quick check

Quiz: Why can a correct requests-per-second plan still fail at a live event? — Request rate is a flow and misses two concurrency-driven costs: memory for merely-open connections, and the fan-out multiplier turning one event into one message per connected client.

Flashcard: Rough memory for 1 million idle TLS WebSocket connections? — 30–60 GB of kernel and user-space buffers at roughly 30–60 KB each, before any application state.