Stacked Resilience Decorator
also called Retry Multiplication, Layered Retry Stack
An anti-pattern in which several independently reasonable layers each add retries or timeouts, so the per-request totals multiply and a slow dependency receives an order of magnitude more load.
A dependency slows from 40 ms to 900 ms. Two minutes later the service is saturated and the dependency is receiving fourteen times its normal request rate. Nothing errored at first, and every line of the code that produced this was reviewed and approved.
Three layers each retry three times: the HTTP client's default, a RetryingRepository decorator added during a flaky integration, and the message consumer's at-least-once redelivery. 3 × 3 × 3 = 27 attempts for one logical request. At healthy latency the first attempt almost always succeeds, so the multiplication is invisible until the moment it is harmful.
Why it matters
Each layer is locally correct, individually reviewable and unaware of the others, so no reviewer sees the product. The failure therefore cannot be found in a code review of any single change, which is what makes it endemic rather than careless.
The second multiplier is concurrency. Every in-flight attempt holds a worker for the duration of its timeout, and because each timeout was chosen locally the outer one usually exceeds the sum of the inner ones. The pool drains and the service stops serving traffic that has nothing to do with the slow dependency. A dependency handling 3% of requests takes down 100% of them.
Implementation patterns
- Retry at exactly one layer and make the others fail fast. The right layer is the one that knows the request's business meaning and its deadline, usually the outermost.
- Use a retry budget rather than a per-call count: cap retries as a share of request volume, for example 10%, so amplification is impossible by construction because the limit is global rather than per call site.
- Propagate an absolute deadline so each layer works with the remaining budget and refuses work it cannot finish, which prevents an inner retry outliving the caller's patience.
- Trip breakers on slow-call rate, not only on error rate. A breaker keyed on errors never opens for a dependency that is merely slow, which is the failure the breaker exists to prevent.
- Alert on attempts per inbound request, per dependency. Above roughly 1.1 in steady state is a design problem waiting for a bad afternoon, and this is the signal that names the fault in seconds.
- Declare the total in one place. A single resilience configuration per outbound dependency, asserted in a test, so the worst-case attempt count is a reviewed number rather than an emergent one.
Industry example
Google's SRE book (2016) describes the same amplification in client retry behaviour and prescribes a server-side retry budget as the control, noting that retries consume the same capacity as first attempts. The mechanism also appears in published postmortems of cascading failures across the industry, where the trigger is a latency increase rather than an outage and the amplifier is client behaviour. The reason it recurs is architectural rather than cultural: any cross-cutting behaviour that can be composed — retries, caches, timeouts, breakers, rate limiters — will be composed by different people at different times.
Failure scenarios
- Load amplification at the worst moment, an order of magnitude more traffic to a dependency that is already struggling.
- Worker-pool exhaustion, so unrelated endpoints fail and the blast radius becomes the whole service.
- Retry storms after recovery, where a queue of pending retries lands at once and knocks the dependency over again.
- Duplicate side effects, where 27 attempts at a non-idempotent operation produce several charges, messages or records.
- Breakers that never open, because every layer measures error rate and the dependency is slow rather than failing.
- Timeout inversion, where an inner timeout exceeds the outer deadline, so the caller has already given up on work the system is still doing.
Trade-offs
Collapsing to one retry layer costs you local robustness: a transient connection reset that an inner immediate retry would have absorbed invisibly now surfaces as a request failure, and some genuinely flaky integrations get worse before the budget is tuned. Centralising the configuration also costs autonomy — a team can no longer add a retry to solve their own problem without changing a shared setting. What you buy is a bounded worst case and a system whose behaviour under latency is predictable rather than emergent.
When not to use it
Retrying at more than one layer is defensible when the layers protect genuinely different failures and the inner one is tightly bounded: a single immediate reconnect inside a connection pool under a 50 ms cap, beneath a business-level retry with a 2-second budget. The test is whether the worst-case total is bounded and written down. If nobody can state the maximum number of attempts one request can cause, the stack is unsafe regardless of how reasonable each layer looks in isolation.
Interview question
Q: A service becomes unavailable whenever one of its six dependencies gets slow, even dependencies used by a small fraction of requests. The team has added circuit breakers at every call site and it has not helped. Walk me through your diagnosis and the first three changes you would make.
What a strong answer covers: asking for attempts per inbound request before anything else · identifying multiplication across layers and the thread-occupancy mechanism as two separate amplifiers · explaining why error-rate breakers do not fire for slow dependencies and what slow-call-rate triggers change · retry budgets over per-call counts, and deadline propagation so inner work cannot outlive the caller · bulkheads or separate pools so one dependency cannot consume all workers · and the governance point that the total must be declared and tested, not emergent.
Quick check
Quiz: Three layers each retry three times. What is the worst-case attempt count for one request, and why is reducing one layer to two attempts an inadequate fix? — 27; because the layers multiply, so the structure still amplifies and only a global budget or a single retry point bounds it.
Flashcard: Which breaker trigger protects against a dependency that is slow rather than failing? — Slow-call rate, the share of calls exceeding a latency threshold; an error-rate breaker stays closed while caller threads pile up.