A live scoreboard JSON endpoint is cached at the CDN with a 5-second TTL and polled every 5 seconds by each viewer. Peak concurrency for an Indian cricket final is in the range Hotstar publicly reported for the 2019 World Cup semi-final - 25.3 million concurrent viewers. Product wants the payload to include the viewer's own fantasy-team score. Roughly what request rate reaches the origin before and after that change?
Show the full answer Hide the answer
The assumptions, stated
25 million clients polling every 5 seconds is 5 million requests per second arriving at the edge. That number is fixed by the client, not by the cache.
What reaches the origin depends on one thing: how many independent cache fills happen per TTL window. Assume the CDN serves this traffic from roughly 40 points of presence with roughly 10 cache machines each — about 400 independent caches. Origin response time 100 ms.
The arithmetic
Shared payload, 5-second TTL, request collapsing on: each cache fetches the object once per TTL. 400 fills / 5 s = roughly 80 requests per second at the origin, whatever the audience does. A public live score is a cheap problem because the origin load is a function of the cache topology rather than of the crowd.
Shared payload, collapsing off: every request arriving while a fill is in flight also goes to the origin. Per cache the arrival rate is 5M / 400 = 12,500 rps, and the in-flight window is 100 ms, so each refresh leaks about 1,250 requests. 400 caches × 1,250 / 5 s = roughly 100,000 rps. Three orders of magnitude, from one configuration flag.
With the fantasy score in the payload: the cache key now varies per viewer. There are 25 million distinct keys, each read once per 5 seconds by exactly one client, so the hit ratio is approximately zero and the origin sees the full 5 million rps. No origin serves that, so the first thing to fail is the origin's connection pool, several minutes before anyone thinks to blame a new field in a JSON body.
Which assumption dominates the error
Not the audience. The count of independent cache fills — POPs × machines × cache tiers — which can be off by 5× between providers and configurations. Enabling a tiered or shield layer collapses all 400 fills into one per region, taking the shared case from roughly 80 rps to single digits. Measure it rather than assuming it: the origin's own request log by cache-fill user agent gives the real number in an afternoon.
What the number rules in or out
- Keep one public document on a short TTL and fetch personal values separately on a longer interval, or push them over the existing connection. Two requests beat one uncacheable request by three orders of magnitude.
- Polling is not the problem. A 5-second TTL plus a 5-second poll means worst-case staleness near 10 seconds, which is wrong for a wicket, so the shared document is the thing to push — not the personal one.
- Cap cache-key cardinality explicitly. Choose a segment rather than a person when personalisation at the edge is unavoidable: 20 segments is a 20-key fan-out and a 95% hit ratio, 25 million is a bypass.
When this is the wrong answer
When concurrency is in the thousands rather than the millions, the personalised payload costs a few hundred rps and a single cached database query answers it. Splitting the response then buys nothing and costs a second request on every tick. The arithmetic above only bites above roughly 100,000 concurrent clients.
Common weak answers
- "Lower the TTL for freshness." Freshness comes from the push path. Halving the TTL doubles origin fills.
- "Scale the origin." 5 million rps of JSON is a CDN's job; buying it at the origin is the most expensive possible answer.
- "Use stale-while-revalidate." It smooths the fill, and does nothing about 25 million distinct keys.