practice

Full-Link Stress Testing

also called Production Stress Test, Shadow-Table Load Test

Driving synthetic load through the live environment with every test request tagged and its writes diverted to shadow tables, so a rehearsal exercises the real topology instead of a staging approximation.

alibabastress-testingproduction-testingshadow-tablestraffic-tagging

A retailer load tests at 5× peak in staging, passes, and falls over at 1.2× peak on the day. The post-incident finding is almost always the same class of thing: a quota shared with a batch job, a connection pool shared by two services nobody tested together, a cache whose real hit rate is 20 points below staging's.

Those defects live in the topology, not in the code, so a test that does not run on the real topology cannot find them. Full-link stress testing runs the rehearsal in production with two safety primitives: every synthetic request carries a tag that propagates through the whole call graph, and every write from tagged traffic lands in a shadow table beside the real one.

Why it matters

A staging environment is a model whose errors are systematic rather than random: the working set is smaller, so hit rates are optimistic; the data is synthetic, so skew and hot keys are absent; shared dependencies are stubbed, so their limits are invisible. The error always flatters, which is why "we passed at 5× and failed at 1.2×" is a recurring story rather than bad luck.

It also rehearses the organisation: dashboards, shedding switches, escalation path and people are exercised alongside the code, which is the half of peak readiness that capacity arithmetic never covers.

Implementation patterns

  • A tag that survives every hop: HTTP header, RPC metadata, message-queue header, and the hand-off into a thread pool where context is most often dropped.
  • A shadow destination for every write and every side effect. Databases are the easy part; email, SMS, payment calls, search indexing, outbound webhooks and analytics each need a tagged path or an explicit block.
  • Metrics and logs split by tag, or the rehearsal turns your own SLO dashboards red.
  • A kill switch that stops the test within seconds, with a rehearsed decision about who pulls it.
  • A ramp with pass criteria agreed in advance: breaking point, failure mode, recovery time.

Industry example

Alibaba Cloud's engineering write-ups describe this practice for Double 11, dating it to 2013: stress tests are performed in the online environment, test requests carry a label that the middleware recognises and that its EagleEye tracing layer propagates through the call graph, and tagged writes are directed into shadow tables beside the real ones so that test data never enters business tables or reported statistics. The driving constraint is a peak that arrives at a fixed instant, an order of magnitude above normal traffic, on a platform too large for a faithful staging copy.

Failure scenarios

  • A dropped tag writes test orders into a real table — the failure the design exists to prevent, and it recurs with every new service and async hop.
  • Side effects nobody enumerated, so the rehearsal sends real emails or calls a real payment provider. The first such event costs more than the programme saves.
  • Shadow tables that drift from the real schema, so the rehearsal exercises an older write path and passes for the wrong reason.

Trade-offs

What it buys: real cache behaviour, real shared-dependency limits, real operational readiness. What it pays: a permanent tax on every team, because tag propagation and shadow destinations become a requirement of every new store, protocol and side effect; a production environment that is also a test environment, which changes the conversation with auditors; and the risk that one missing propagation corrupts real data.

This is insurance against a specific, dated, unreschedulable event. With no such event, the premium buys nothing.

When not to use it

If the peak is 2× rather than 20×, do not build this. Realistic staging at production data volumes, plus a read-only dark-traffic test in production — mirror real requests to a shadow fleet and discard the responses — captures most of the value and writes nothing.

If you cannot enumerate your side effects, do not start. The maturity ladder is realistic staging → production read-only replay → tagged write isolation on one critical path → full link. Teams that jump to the last rung often spend a year on tag propagation and then find their constraint was a single connection pool they could have found in week one.

Interview question

Q: "We want to stress test in production before our biggest sale. Tell me what you would build first, what you would refuse to do, and how you would stop the test."

What a strong answer covers: tag propagation as the first primitive and the async hand-off as its weak point; an inventory of writes and side effects before any traffic is generated, with the irreversible ones blocked rather than shadowed; tag-aware metrics so the test does not blind the on-call; a kill switch as a control with a named owner; and the judgement to propose read-only production replay instead when the peak does not justify the permanent tax.

Quick check

Quiz: What two primitives make a production stress test safe? An end-to-end propagated tag on every test request and a shadow write destination for every data store and side effect.

Flashcard: Why does a staging load test fail in a consistent direction rather than randomly? — Its working set, data skew and shared dependencies are all smaller or absent, and every one of those errors makes the system look faster than it is.