A live-sport streaming service is planning for a cricket final. Marketing says "be ready for a big audience". The semi-final peaked at 1.8 million concurrent viewers and climbed from 400000 to 1.8 million in roughly 90 seconds at the first ball. Segments are 4 seconds long. Roughly what does the final demand, and which of those numbers actually constrains the design?
Show the full answer Hide the answer
The assumptions, stated
A final draws more than a semi-final, and the multiple is the uncertain part. Take 3x to 5x the semi-final peak and carry the range: roughly 5 to 9 million concurrent viewers, call it 7 million as the planning figure. Assume the same arrival shape, because the trigger is the same — the first ball — and assume an average bitrate of about 3 Mbps across a mixed device population. The only figures here measured in production are the semi-final's peak and its 90-second ramp; everything else is derived from them and should be labelled as such in the plan.
The arithmetic
- Manifest and segment requests. Each viewer fetches one segment every 4 seconds, so request rate is viewers divided by segment duration. 7 million / 4 s = about 1.75 million requests per second to the edge, plus a similar rate of manifest refreshes on many players.
- Egress. 7 million x 3 Mbps = about 21 Tbps at peak.
- The ramp. The semi-final added 1.4 million viewers in 90 seconds, or about 15500 new sessions per second. Scale that by the same 4x and the final adds roughly 60000 sessions per second during the opening minute.
The number that constrains the design
Not the peak. The ramp rate. An instance that takes 60 to 120 seconds to boot, warm its caches and pass a health check cannot join in time to serve a surge that completes in 90 seconds, so reactive autoscaling arrives after the event it was meant to absorb. Every layer with a warm-up longer than the ramp has to be provisioned before the toss, not scaled during the over.
Second-order consequences fall out of the same number. 60000 new sessions per second is 60000 TLS handshakes per second and 60000 authorisation checks per second against whatever holds entitlements, and those are the components that see the surge undiluted — a cache cannot help a request that has never been made before.
Which assumption dominates the error
The 3x-to-5x multiple, by a wide margin. Everything else is arithmetic on measured quantities. That is why the planning output is a pre-provisioned floor sized at 5x with a documented trigger to extend to 9x, rather than a single number: the uncertainty lives in one term and belongs in the plan rather than hidden in it.
What the number rules in and out
It rules out reactive autoscaling on the entitlement and handshake path, rules out any cache that must be populated by live traffic, and rules in pre-warmed capacity, a queue in front of authorisation, and a degradation ladder agreed in advance — drop to a single bitrate ladder, defer personalisation, serve a stale entitlement decision for 60 seconds. It also rules in the rehearsal: the plan is worth nothing if 7 million sessions of load have never been applied to it.
When this is the wrong analysis
For a system whose demand grows over hours rather than seconds, the ramp rate is irrelevant and pre-provisioning to peak costs real money for nothing. The test is a ratio: divide the time to add capacity by the time the surge takes to complete. Below about 0.1, prefer autoscaling and skip this exercise entirely; above 1, no amount of autoscaling configuration will save the event and the money goes into idle capacity instead. The uncomfortable part of that trade is that the pre-provisioned floor is paid for whether or not the audience arrives, which is why the 5x figure and the trigger to extend it both belong to the business rather than to the platform team.