advanced 2 min answer

Alibaba reported a peak of 583000 orders per second during Singles' Day 2020, far above its ordinary daily rate. A retailer with a 40x annual peak asks whether it should provision for peak all year, since that is what the big platforms clearly do. What did a platform in that position actually buy, and where does copying it go wrong?

alibabapeak-capacitydegradationrehearsalflash-sale
Show the full answer Hide the answer

The situation they were in

A single scheduled event, known to the second, carrying a large share of the year's revenue. That combination is unusual and it is what makes their approach rational: a peak you can put in a calendar can be rehearsed, and revenue concentrated into hours can justify capacity that idles for the rest of the year.

What they chose

Capacity was one part of it and not the interesting part. Alibaba Cloud's own write-ups describe full-link stress testing: a rehearsal run against the production environment with synthetic traffic tagged so it can be told apart from real business traffic, plus pre-configured traffic limiting and degradation that can be adjusted during the event. On the data side, Alibaba built its own distributed relational database, OceanBase, used behind Alipay, because the write peak on the transaction path was the constraint off-the-shelf sizing could not answer.

Why it fit their constraints

At that order rate, the binding limits are contention and coordination on a small number of hot rows, not aggregate CPU. Adding servers does not reduce contention on one product's stock counter. So the engineering went into the paths that contend, and the rest of the system was given a documented way to get out of the way.

What it cost them

A permanent capacity organisation, rehearsals that consume engineering time for weeks, and a codebase where every feature carries a degraded mode that has to be maintained and tested. That is a standing tax paid all year for a handful of hours.

Where copying it goes wrong

Run the arithmetic for the 40x retailer. Provisioning for peak all year multiplies steady-state compute by about 40 to serve, say, four hours of genuine peak, which is roughly 0.05% of the year at full utilisation. The same money buys a rehearsal programme, a load-shedding path and a static browse fallback, and those work for unscheduled peaks too — which matters, because an unscheduled peak is the case the retailer actually has.

The transferable moves, in order:

  1. Shed rather than scale on the paths that can wait. A published queue position during a sale is better than a timeout and far cheaper than 40x capacity.
  2. Keep the reservation path correct and small. Everything else may serve stale or degrade.
  3. Rehearse against production, not staging, with the synthetic traffic tagged so it can be excluded from business metrics. That is the copyable part, at any scale.
  4. Pre-declare what gets switched off, so the choice is made in a calm room rather than during the event.

When this is the wrong answer

If the peak-to-trough ratio is 3x rather than 40x, all of this is over-engineering: buy the headroom, keep the architecture simple, and spend the saved complexity elsewhere. The ratio and whether the peak is scheduled are the two facts that decide. Building degradation modes and rehearsal programmes for a 3x peak is the mirror image of provisioning 40x all year: both pay for a problem the business does not have.