practice

Full-Link Stress Test

also called Full-Chain Load Test, Production Rehearsal

A load test generated against the production environment across the entire request path including shared dependencies and third parties, isolated by a propagated traffic marker so synthetic writes never touch real data.

load testingtesting in productionalibabasingles daycapacityshadow data

An annual peak many multiples of ordinary traffic, on a known date, across hundreds of services. The result is binary and public. A staging load test returns a reassuring number, and the number is about a different system.

That is the problem full-link stress testing exists to solve: not "is this service fast enough" but "will this system, as configured today, with its real shared dependencies and real partners, survive the peak." No environment other than production can be asked that question.

The practice is load generation against production, with the entire path exercised end to end, and with isolation machinery that keeps synthetic traffic from corrupting real state.

Why it matters

The gap between staging and production is not size. It is four things, and each produces a different class of surprise:

  • Configuration divergence. Pool sizes, timeouts, feature flags, rate limits, DNS TTLs, kernel settings. Staging drifts continuously, and the divergences that matter are the undocumented ones.
  • Shared dependencies are shared. Production services contend for the same database, cache and network path as everything else. A service tested in isolation is a different system.
  • Third parties. Payment providers, carriers, identity services — a first-order risk, represented in staging by a stub answering in 5 ms.
  • Data shape. A join that is fine over 100,000 rows picks a different plan over 4 billion, and the change is a cliff rather than a slope.

Implementation patterns

The machinery must be built in this order, because each piece depends on the one before, and because steps 1–3 are what make step 5 safe:

  1. A traffic marker that propagates everywhere — a header, context value or message attribute, carried through every synchronous call, every queue message and every async job. The places it is forgotten are exactly where the damage happens.
  2. Data isolation keyed on the marker. Marked writes go to shadow tables or shadow key prefixes. Reads may come from real data; writes must not touch it.
  3. Side-effect suppression at a single egress boundary. No emails, SMS, push notifications, card charges, partner webhooks or ledger entries. Enforced in one layer that drops marked traffic, never by asking each service to remember — this is the piece with legal consequences.
  4. The marker as a telemetry dimension, so real-user impact can be read separately from synthetic load while the test runs.
  5. A one-step abort with a measured time-to-stop, operable by whoever is watching.
  6. Realistic traffic composition, not uniform load: the real request mix in real proportions, including the long tail. A flat test of the cheapest endpoint proves nothing about the peak.

Industry example

Alibaba has published full-link stress testing as its Singles' Day rehearsal: load generated in the production environment, with traffic shading and data isolation so synthetic requests are kept away from real records, run with partner organisations in scope. The published accounts describe the practice maturing over several annual cycles, with the number of rehearsal sessions needed falling as the platform improved — which is the useful detail. The capability was built incrementally into the platform, not assembled before one event. Marker propagation across hundreds of services is the kind of change that cannot be retrofitted in a quarter; it has to live in the service template.

Failure scenarios

  • Unsuppressed side effects. Real emails to real customers, real webhooks to partners, or a write to a financial ledger. The characteristic disaster, and it comes from running load before step 3 is complete.
  • A marker that drops at one hop. Usually an async boundary — a queue consumer that does not read the attribute, a scheduled job spawned by a marked request. Synthetic writes then land in real tables, and nobody notices until a reconciliation days later.
  • Degrading real users. The load is real load on real infrastructure. This is why rehearsals run in low-traffic windows with a live abort and people watching.
  • Shadow tables with different statistics. If shadow tables are empty while real ones hold billions of rows, the query plans under test are not the production plans, and the test's central claim is void.
  • A stale rehearsal. The system changes after the test. A single rehearsal weeks before the peak proves something about a system that no longer exists.

Trade-offs

Choose Gains Pays
Full-link production rehearsal The only answer about the real system, including shared deps and partners Platform-wide instrumentation, standing data-corruption risk, real user exposure
Staging load test at production scale No corruption risk, repeatable, safe to run often Tests a system whose configuration and dependencies differ in undocumented ways
Per-service load test against production-sized data Cheap, finds query-plan and resource ceilings Says nothing about contention between services or about partners

When not to use it

Build it only if the peak is both large and predictable. The practice exists to de-risk a known, enormous, dated event. A business with steady traffic and 20% annual growth gets equivalent confidence from load-testing one service at a time against production-sized data, plus autoscaling with genuine headroom and a tested load-shedding path — at a fraction of the engineering cost and none of the data-corruption risk.

The sequencing rule matters more than the ambition: a team that generates shadow traffic before the isolation machinery is complete and independently verified produces an incident, not a rehearsal. This is the most predictable failure in the area. If the honest answer to "is the marker honoured at every async boundary?" is "probably", the answer to "should we run load today?" is no.

And if the real question is "what breaks first under load", a load test against one service plus a dependency map answers it for far less, because the first thing to break is usually a single resource ceiling rather than an emergent whole-system effect.

Interview question

Q: Leadership wants a production load test before the festive peak, in six weeks. The platform has no traffic marker. What do you tell them?

What a strong answer covers: that marker propagation and side-effect suppression are platform-wide changes that cannot be delivered safely in six weeks, said plainly rather than hedged · the risk of proceeding, named concretely (partner webhooks, ledger writes, emails) · the alternative that fits the window: per-service load tests against production-sized data, a dependency map, verified autoscaling headroom and a tested shedding path · building the marker into the service template for next year · and being specific about what each option does and does not prove.

Quick check

Quiz: Which piece of full-link isolation machinery has legal consequences if it is incomplete, and where is it enforced? — Side-effect suppression, enforced at a single egress boundary that drops marked traffic rather than in each service.

Flashcard: Why can a staging load test not answer whether the system will survive the peak? — Configuration drifts undocumented, shared dependencies are not shared, third parties are stubs, and query plans change discontinuously with real data volume.