Alibaba reported a peak of 583,000 order-creation requests per second during the 2020 Singles' Day festival on 11 November. Demand of that shape is produced by a countdown every client sees at the same instant. What does scheduled demand change about the mobile client's responsibilities, and where would copying the approach be a mistake?
Show the full answer Hide the answer
The situation that creates the number
583,000 orders per second is the documented peak (reported for Singles' Day 2020). What makes it an architecture problem rather than a capacity problem is that the arrival time is chosen by a clock rather than by demand. Organic traffic has a peak-to-mean ratio of three to five and arrives with natural jitter. A countdown removes the jitter: tens of millions of clients act within the same second because they were told to.
The mechanisms below are the general shape of a client designed for scheduled demand. They are not a claim about any internal system at Alibaba, which has not been described in the sources used here.
What moves onto the client
- Pre-warming before the gun. Assets, cart contents, delivery address and a payment authorisation token are fetched in the minutes before the event, so the request at T plus zero carries one write and no reads. This is the single largest reduction in peak work, and it trades a predictable earlier load for an unsurvivable later one.
- Client-side admission. The server issues a start token with a randomised offset, and the client waits it out. A three-second jitter window spreads a one-second spike across three, cutting the instantaneous rate by roughly two thirds without changing anything server-side.
- An idempotency key created on the device and written to disk before the first attempt. On a congested network the response is lost far more often than the request, so the retry is certain. Without a device-generated key, the retry is a second order.
- Server-directed retry. The client must honour a
Retry-Afterinstead of running its own backoff loop, because every client's own loop is synchronised with every other client's. - No trust in the device clock. The countdown runs on a server time offset applied to a monotonic clock, or a phone that is two minutes fast becomes a denial-of-service client.
What it costs
Everything in the window is rehearsed and frozen: capacity pre-provisioned rather than autoscaled, a code freeze, a feature set deliberately reduced, and a client release shipped weeks early because the app store is not under your control and old versions never disappear. Pre-warming is a real bill of its own: 2 MB of assets to 50 million clients is about 100 TB of egress spent before a single order exists.
Where copying this is a mistake
A platform whose peak is organic does not need any of it. At three to five times the daily mean, autoscaling, a queue and a sensible timeout budget are sufficient, and client admission tokens add a protocol, a failure mode and a release dependency for no benefit. The flip condition is specific and easy to test: does your traffic arrive because a clock told it to? Ticket sales, exam results, payroll days and game launches qualify. A retail checkout does not.
Common weak answers
- "Autoscale harder." Scaling reacts over tens of seconds to minutes. The event is decided in the first second, before any new capacity exists.
- "Put a queue in front of it." Necessary and insufficient: a queue that accepts 50 million enqueues in one second still needs the admission decision to be made before the request leaves the phone.
- "Rate-limit at the edge." It protects the origin and converts a spike into millions of rejected customers, which is the outcome the design exists to prevent.