advanced 4 min answer

JioStar reported a peak of 55.2 million concurrent viewers on JioHotstar for the 2025 IPL final. Take a platform of that shape. At the last ball, product wants a "match won" push sent to 120 million installs carrying a deep link into the highlights screen. Video is on a CDN and the API tier is sized for 1.5x the match-time peak. What happens in the first 90 seconds after the send begins and what stops it?

hotstarpush-notificationsthundering-herdfan-outpeak-event-readiness
Show the full answer Hide the answer

Second by second

Seconds 0–20. The fan-out service starts walking a 120-million-row token table. Submission rate is bounded by your own sender fleet and its multiplexed connections to the push providers, not by the providers' appetite, so the practical send rate is whatever your fan-out tier can sustain. At 500,000 sends per second the blast takes four minutes; at 50,000 it takes 40 minutes and the product requirement is already lost.

Seconds 5–45. Devices begin receiving. Every tap opens the app, and a cold app open is not one request. It is a token refresh, a session validation, a config fetch, a personalised home or highlights payload, and usually an advertising call: eight to fifteen requests, several of them uncacheable. One million opens in ten seconds is therefore on the order of one million API requests per second, against a tier sized for 1.5x the match-time peak.

Seconds 20–90. The API tier saturates, p99 crosses the client timeout, and the clients retry. Retry turns one failed open into two or three requests, so offered load rises as capacity falls. Video keeps playing, because the CDN absorbed it, which makes the dashboard confusing: bytes delivered look healthy while every logged-in action fails.

Where it amplifies

  • The trigger is a single instant for everyone. Unlike organic traffic, a push has no arrival distribution except the one you impose.
  • Cold starts on the client. Apps that have been backgrounded for an hour re-handshake TLS, re-resolve DNS and re-authenticate. The per-open cost is at its maximum exactly when concurrency is too.
  • Dead tokens still cost. A push target list grows faster than the installed base, so a large share of the fan-out spend reaches devices that no longer exist, while your delivery percentage describes the list rather than an audience.
  • Queued pushes arrive in bursts. Devices that were offline receive their backlog on reconnect, so the stampede has a long second wave hours later.

What the user sees

A notification, then a spinner, then a stale screen. The push converted a successful broadcast into a failed product moment, and the failure is concentrated in the users who were not already watching, which is the audience the push existed to reach.

What stops it

  1. Spread the send and say so in the plan. Cohort the token list and release it over several minutes with jitter. This costs timeliness, which is a product decision, not an engineering one: a "match won" push is still correct 4 minutes later.
  2. Make the landing screen free. If the deep link lands on a CDN-cacheable highlights page with stale-while-revalidate, an app open costs zero origin calls. This is the single largest lever, because it changes the per-open request count rather than the open rate.
  3. Collapse per device. A collapse identifier on the push means a device that was off receives one notification, not six.
  4. Admission control with priority classes at the API edge. Authentication and playback entitlement are admitted; recommendations, badges and advertising calls are shed first. Without classes, the shed is random and takes out logins.
  5. Pre-scale against the schedule. Sports demand is scheduled, which means capacity can be in place before the event rather than discovered during it. Scheduled demand is the one peak autoscaling does not need to guess at.

What would have to be true for this to self-heal

Only one thing: that an app open requires no origin call. A stampede into cacheable content decays on its own as the CDN fills. A stampede into a personalised screen does not self-heal, because every retry is a cache miss by construction. Treat the question "is the landing screen personalised" as a capacity decision, because that is what it is.

Common weak answers

  • "Autoscale the API tier." Scheduled instances take minutes to become useful and the event lasts 90 seconds. Autoscaling is for the hour after the stampede, not for the stampede.
  • "Trust the provider's delivery figure." It describes submissions to a token list, not arrivals at an audience, and a list inflated by dead tokens makes the number look better as it gets worse.
  • "Rate limit the clients." Shedding the app opens you asked for converts a capacity problem into a visibly broken product, unless the shed is by priority class and the authentication path is protected.
  • "Send it from the fan-out service as fast as it can go." The send rate is the one variable you fully control, and spending it all at once is the choice that creates the incident.