Alibaba has published that it rehearses Singles' Day with full-link stress testing — load generated against the production environment, using traffic shading and data isolation. Your platform faces a comparable annual peak. What does testing in production buy that a staging load test cannot, and what has to be built first?
Show the full answer Hide the answer
The situation they were in
An annual peak that is many multiples of ordinary traffic, falling on a known date, across a system that spans hundreds of services and external partners. The result is binary and public: it works, or it fails in front of the entire market. There is no gradual ramp during which problems surface gently.
Alibaba's published approach is a rehearsal in the production environment — not a copy of it — with mechanisms to keep the synthetic traffic from corrupting real data.
What a staging load test cannot tell you
This is the crux, and it is not about staging being "smaller".
- The configuration is different, and nobody knows how. Connection-pool sizes, timeouts, feature flags, rate limits, DNS TTLs, instance types, kernel settings. A staging environment diverges continuously, and the divergences that matter are precisely the ones nobody documented.
- The shared dependencies are not shared. Production services contend for the same database, the same cache cluster, the same network path as everything else running in production. Staging tests a service in isolation and therefore tests a different system.
- Third parties are not in staging. Payment providers, carriers, identity providers, partner APIs. Their rate limits and their behaviour under your peak are a first-order concern and are usually represented by a stub that always answers in 5 ms.
- The data is not the same shape. Query plans depend on table sizes and value distributions. A join that is fine over 100,000 staging rows picks a different plan over 4 billion production rows, and the difference does not scale linearly — it is a cliff.
The question a production rehearsal answers is "will this system survive", and no other environment can be asked that question.
What has to be built first
In order, because each depends on the last:
- A traffic marker that propagates everywhere. A flag on the request — header, context, message attribute — carried through every synchronous call, every queue message, every async job. This is the foundation, and the places it is forgotten are exactly where the damage happens.
- Data isolation keyed on that marker. Writes from marked traffic go to shadow tables or shadow key prefixes rather than the real ones. Reads may come from real data; writes must not touch it.
- Side-effect suppression. No emails, no SMS, no push notifications, no real card charges, no webhooks to partners, no entries in financial ledgers. This is the part with legal and reputational consequences, so it is enforced at the boundary — a single egress layer that drops marked traffic — rather than by each service remembering.
- Marked traffic in observability. Every metric and log dimensioned by the marker, so real-user impact can be read separately from synthetic load during the test.
- An abort switch that works in one step. A single flag that stops generation and drains, operable by whoever is watching, with a tested time-to-stop.
What it cost them
Instrumenting every service to propagate and honour the marker is a platform-wide change, which is exactly the kind of work that is impossible to retrofit in a quarter and must be built into the service template. The published accounts describe this maturing over several annual cycles rather than arriving at once, with the number of rehearsal sessions falling as the platform and the practice improved.
There is also a standing risk that cannot be fully removed: the marked traffic is real load on real infrastructure. A rehearsal can degrade the production experience of real users, which is why it is run in a low-traffic window with a live abort.
When copying this would be the wrong answer
Build it only if your peak is both large and predictable. Full-link production stress testing exists to de-risk a known, enormous, dated event. A business with steady traffic and 20% annual growth gets the same confidence from load-testing a single service at a time against production-sized data, plus autoscaling with real headroom — at a fraction of the engineering cost and none of the data-corruption risk.
The sequencing matters more than the ambition. A team that runs shadow traffic before the side-effect suppression is complete will send real emails, or worse, write to a real ledger. The pattern is not "load test in production"; it is "build five pieces of isolation machinery, verify each independently, and only then generate load". Teams that invert this order produce an incident rather than a rehearsal, and it is the most predictable failure in this entire area.