advanced 3 min answer

Disney+ Hotstar reported a peak of 59 million concurrent viewers during the 2023 cricket World Cup final, having passed 53 million at the semifinal. You are sizing the real-time engagement features for an event in that class. Estimate what the concurrency implies, and say which number decides the architecture.

hotstarconcurrencycapacity planninglive eventsfan-outestimation
Show the full answer Hide the answer

The assumptions, stated

Video delivery is a CDN problem and largely solved by buying capacity. The hard part is everything stateful that sits beside the stream, so the estimate is about connections, fan-out and the arrival shape.

  • 59M concurrent viewers at peak, reported by Disney for the 2023 final.
  • Persistent connections for live features for perhaps 40% of viewers — 24M concurrent connections.
  • Engagement events (reactions, poll votes) from 1% of viewers in a normal minute, spiking to 10% within 10 seconds of a wicket.
  • Fan-out of a scoreboard update to every connected client.

The arithmetic

Connection memory. An idle TLS WebSocket costs roughly 30–60 KB of kernel and user-space buffers. Take 45 KB:

24,000,000 × 45 KB ≈ 1.1 TB of connection memory, before application state

At perhaps 500,000 connections per well-tuned node, that is about 50 nodes just to hold the sockets — which is affordable, and tells you the connection tier must be separate from anything that does work, because it has to be scaled on a different axis.

Inbound event rate at a spike. 10% of 59M in 10 seconds:

5,900,000 events ÷ 10 s = 590,000 events/s arriving

Outbound fan-out, the number that decides everything. One scoreboard update to 24M connections:

24,000,000 messages per update

At one update per second that is 24M messages/s. Three orders of magnitude above the inbound rate. This is the whole design.

Which assumption dominates the error

Not the viewer count — that is reported. It is the fan-out frequency, and it is the one under your control. Updating a scoreboard once per second rather than once per five seconds is a 5x difference in the dominant cost, and no user can perceive it. The second most sensitive assumption is the connection fraction: 40% versus 70% nearly doubles the connection tier.

What the number rules in and out

  • Ruled out: per-client computation on fan-out. At 24M messages per update there is no budget for per-user personalisation. The broadcast payload must be identical for everyone, built once, and pushed. Personalisation happens on the client or not at all.
  • Ruled in: aggressive batching and coalescing. Clients receive at most one update per interval, carrying the merged state. Coalescing is the only lever with the right exponent, because it divides the dominant term directly.
  • Ruled in: a connection tier that does nothing but hold connections and relay, scaled independently, with no shared state, so it can be sized for 1.1 TB of buffers without dragging logic with it.
  • Ruled out: sizing from the average. The arrival shape is the problem. 590,000 events/s arriving in 10 seconds after a wicket is not a traffic level, it is a step function, and autoscaling reacting in 60–90 seconds is irrelevant to it. Capacity must be pre-provisioned to the predicted peak before the event starts — which is possible, because the schedule is known months ahead.
  • Ruled in: a tested degradation ladder. At the peak, shed in a defined order: reactions first, then poll results, then comment fan-out, then update frequency, leaving the video and the score. Decide this in advance and rehearse it.

When not to build any of this

At 100,000 concurrent viewers — a figure most live products never exceed — nearly all of it is unnecessary. 40,000 connections is 1.8 GB of buffers, which is a couple of instances, and 40,000 messages/s of fan-out is comfortable for a single managed pub/sub topic with an off-the-shelf WebSocket gateway. The architecture above is justified by the fan-out term crossing roughly a million messages per second, and not before. Building it earlier buys an operational burden and a distributed system to debug, in exchange for headroom you will not use.

The decision rule: compute the fan-out term first, because it is the one that is cubic in nothing and linear in the thing you control. If coalescing can keep it under about a million messages per second, buy a managed service and spend the engineering somewhere else.