Traffic Shading
also called Traffic Marking, Shadow Tagging, Request Tainting
Marking a request as synthetic and propagating that mark through every hop so downstream services can route its writes to shadow storage and suppress its side effects, which is the single mechanism that makes production load testing safe.
A shadow load test writes 40,000 orders at 2,000 per second for 20 minutes. Thirty-nine thousand land in shadow tables. One thousand — 2.5% — land in the real orders table, because one service consumes from a queue and does not read the marker from the message attributes. Nobody notices for four days; the discovery is a reconciliation mismatch, and the fix is identifying and deleting a thousand rows indistinguishable from real orders.
Traffic shading is the mechanism meant to prevent this, and that is its characteristic failure: not the absence of the mechanism, but a single hop that does not honour it.
The pattern is simple. A request entering the system for test purposes is marked; every service reads the mark, acts on it, and passes it on. Since the marker is the only thing distinguishing a synthetic order from a real one, its propagation is a correctness property of the whole estate rather than a feature of one service.
Why it matters
Isolation cannot be achieved at the entry point. A separate test endpoint or load balancer fails as soon as the request fans out, because a service three hops down has no way to know where it came from. The marker carries that knowledge to where the decision must be made, which is at the write.
It also makes the decision local and uniform: each service answers one question and none needs to know about the test, the generator or the topology, which is what lets the pattern scale to hundreds of services.
And it is worth naming as a pattern because it fails silently and partially. A hop that drops the marker produces real writes indistinguishable from real data, with no error and no alert, found days later or never.
Implementation patterns
- One canonical marker in the transport's standard metadata — an HTTP header, gRPC metadata key, message attribute or context value. Never a field in the business payload.
- Propagate in shared middleware, never per service. The library that propagates trace context should carry the marker, for the same reason: a hundred teams will not each remember.
- Default to deny at the storage layer. Derive the shadow-routing decision in the data-access layer from ambient context rather than asking each write path to check. A service that forgets then writes to the shadow table by construction, inverting the failure from silent corruption to a test that writes nowhere useful.
- Suppress side effects at one egress boundary. One outbound gateway for email, SMS, push, webhooks and payments, dropping anything marked: one place to audit and to get right.
- Fail closed on ambiguity. A consumer that cannot tell whether a message is marked must reject it; assuming real is the corrupting choice.
- Test the propagation continuously. A canary synthetic request down every path, asserting at each terminal store that the write landed in shadow. This is the test of the test, and its absence is why the opening scenario happens.
- Dimension telemetry by the marker, so real and synthetic load can be read apart during the run.
Industry example
Alibaba's published description of full-link stress testing names traffic shading and data isolation as the enabling mechanisms for generating Singles' Day rehearsal load against the production environment. The pattern recurs wherever shadow traffic is used — proxy request mirroring, dark launches, shadow deployments — and the common thread is that the hard part is never generating load, it is guaranteeing that a hundred independently-deployed services all honour one bit of metadata. Organisations that run it safely built propagation into their service scaffolding, which makes it a platform capability rather than a test tool.
Failure scenarios
- The dropped hop at an async boundary. Queue consumers, scheduled jobs and retry workers, because the marker must survive serialisation and those paths are hand-written.
- Shadow reads with real side effects. A marked request reads a real customer record and triggers a recommendation update or audit entry against them.
- Cache poisoning. Marked requests populate a shared cache with shadow data, later served to real users. The cache key must include the marker, and frequently does not.
- Counter contamination. Synthetic traffic increments business counters, inventory reservations or rate-limit buckets, distorting the state it was meant to leave untouched.
- A marker anyone can set. If an external caller can set the header, they can probe which paths honour it. Strip it at the edge; set it only from trusted generators.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Marker propagated estate-wide | Production rehearsal becomes possible at all | Platform-wide change; one unhonoured hop corrupts data silently |
| Separate test environment | No corruption risk whatever anyone forgets | A different system: different config, no real partners |
| Read-only mirroring, no writes | Most of the load realism, nothing to isolate | Cannot exercise the write path, usually the bottleneck |
When not to use it
If only the read path matters, mirror requests at a proxy and discard the responses. Realistic read load, no marker to propagate, no isolation to get wrong, and for a read-heavy service it answers most of the capacity question. Reach for shading only when the write path needs testing, which is when it is also the risk.
Do not adopt it partially. A marker honoured by eight services of twelve is worse than none, because the generator is run on the assumption it works. The rule is binary: either every path that writes or produces a side effect honours it and that is continuously verified, or shadow traffic does not run. "Mostly propagated" is the configuration that produces the thousand real orders.
For a small estate — a handful of services, one team — the simpler isolation is a separate schema selected by a deployment-time flag, with a copy of production data. Far less to get wrong, and sufficient where contention between services is not the concern.
Interview question
Q: You are asked to review a shadow-traffic design before its first production run. What do you check, in order?
What a strong answer covers: where the marker is set and whether it is stripped at the edge · tracing it through every async boundary by name, since those are where it drops · whether the storage layer defaults to shadow on ambiguity, so a forgetful service fails safe · side-effect suppression at a single egress boundary rather than spread across services · whether cache keys include the marker · what verifies propagation continuously, treating its absence as blocking · and the abort path with its measured time-to-stop. A strong answer refuses the run if propagation is unverified, and says so.
Quick check
Quiz: Why should the data-access layer default to shadow storage when the marker is ambiguous? — It inverts the failure direction: a service that forgets to check writes harmlessly to shadow rather than silently corrupting real data.
Flashcard: Where does traffic shading usually fail? — At async boundaries. Queue consumers, scheduled jobs and retry workers must carry the marker through serialisation, and those paths are hand-written, so they are where real writes leak in silently.