advanced 3 min answer

A live scoreboard JSON endpoint is cached at the CDN with a 5-second TTL and polled every 5 seconds by each viewer. Peak concurrency for an Indian cricket final is in the range Hotstar publicly reported for the 2019 World Cup semi-final - 25.3 million concurrent viewers. Product wants the payload to include the viewer's own fantasy-team score. Roughly what request rate reaches the origin before and after that change?

cdncache-keyrequest-collapsingpollinghotstar
Show the full answer Hide the answer

The assumptions, stated

25 million clients polling every 5 seconds is 5 million requests per second arriving at the edge. That number is fixed by the client, not by the cache.

What reaches the origin depends on one thing: how many independent cache fills happen per TTL window. Assume the CDN serves this traffic from roughly 40 points of presence with roughly 10 cache machines each — about 400 independent caches. Origin response time 100 ms.

The arithmetic

Shared payload, 5-second TTL, request collapsing on: each cache fetches the object once per TTL. 400 fills / 5 s = roughly 80 requests per second at the origin, whatever the audience does. A public live score is a cheap problem because the origin load is a function of the cache topology rather than of the crowd.

Shared payload, collapsing off: every request arriving while a fill is in flight also goes to the origin. Per cache the arrival rate is 5M / 400 = 12,500 rps, and the in-flight window is 100 ms, so each refresh leaks about 1,250 requests. 400 caches × 1,250 / 5 s = roughly 100,000 rps. Three orders of magnitude, from one configuration flag.

With the fantasy score in the payload: the cache key now varies per viewer. There are 25 million distinct keys, each read once per 5 seconds by exactly one client, so the hit ratio is approximately zero and the origin sees the full 5 million rps. No origin serves that, so the first thing to fail is the origin's connection pool, several minutes before anyone thinks to blame a new field in a JSON body.

Which assumption dominates the error

Not the audience. The count of independent cache fills — POPs × machines × cache tiers — which can be off by 5× between providers and configurations. Enabling a tiered or shield layer collapses all 400 fills into one per region, taking the shared case from roughly 80 rps to single digits. Measure it rather than assuming it: the origin's own request log by cache-fill user agent gives the real number in an afternoon.

What the number rules in or out

  • Keep one public document on a short TTL and fetch personal values separately on a longer interval, or push them over the existing connection. Two requests beat one uncacheable request by three orders of magnitude.
  • Polling is not the problem. A 5-second TTL plus a 5-second poll means worst-case staleness near 10 seconds, which is wrong for a wicket, so the shared document is the thing to push — not the personal one.
  • Cap cache-key cardinality explicitly. Choose a segment rather than a person when personalisation at the edge is unavoidable: 20 segments is a 20-key fan-out and a 95% hit ratio, 25 million is a bypass.

When this is the wrong answer

When concurrency is in the thousands rather than the millions, the personalised payload costs a few hundred rps and a single cached database query answers it. Splitting the response then buys nothing and costs a second request on every tick. The arithmetic above only bites above roughly 100,000 concurrent clients.

Common weak answers

  • "Lower the TTL for freshness." Freshness comes from the push path. Halving the TTL doubles origin fills.
  • "Scale the origin." 5 million rps of JSON is a CDN's job; buying it at the origin is the most expensive possible answer.
  • "Use stale-while-revalidate." It smooths the fill, and does nothing about 25 million distinct keys.