concept

Retry Amplification

also called Retry Storm, Work Amplification

The multiplication of load through a call graph when every layer retries independently, so a brief degradation becomes sustained overload that persists after the original fault is gone.

retriescascading failureload sheddingbackoffanti-patternmetastability

A database slows down for fifteen seconds. Forty minutes later it is still unavailable, the original cause has long since cleared, and nothing anyone does brings it back short of turning traffic off.

The arithmetic explains it. Each layer retries three times, which sounds modest. Through four layers:

3 × 3 × 3 × 3 = 81 attempts for one user request

An 81x load multiplier appearing exactly when the system is least able to serve it. Each layer's policy is individually reasonable, and nobody chose 81.

Retry amplification is that multiplication. It is the commonest engine of a metastable failure — the retries generate the load that causes the failures that generate the retries, so the system does not recover when the trigger clears — and the distinct thing worth knowing about the amplification itself is that the multiplier is a design parameter nobody set, and there are specific controls that bound it.

Why it matters

It is invisible in any single service's design review, because every individual retry policy looks prudent. The multiplication exists only in the composition, so no team owns it and no review catches it. That is the distinguishing feature: this is an emergent property of a call graph, not a defect in a component.

The second point is the one most often confused. Exponential backoff with jitter spreads retries in time; it does not reduce their number. Against a dependency that is down rather than briefly blipping, 81 attempts spread over 30 seconds is still 81 attempts. Backoff addresses the synchronised wavefront; only a budget addresses the multiplier.

Implementation patterns

  • Retry at one layer only, and make it the one closest to the fault. The single most effective change. Everything above it fails fast and propagates. This turns 81 into 3 and is usually a policy decision rather than an engineering project.
  • Retry budgets, not retry counts. Cap retries as a fraction of total requests — "retries may not exceed 10% of traffic to this dependency" — and when the budget is exhausted, fail without retrying. This is the mechanism that bounds amplification by construction, because it is a global limit rather than a per-request one. Envoy and gRPC both support it.
  • Deadline propagation. Pass the caller's remaining time budget with the request; if a retry cannot complete inside it, do not attempt it. Retrying past the caller's deadline is pure amplification — the work is guaranteed to be discarded.
  • Circuit breakers that trip on slow calls, not only errors. A slow dependency produces no errors, so an error-rate breaker stays closed while threads pile up — precisely the condition that feeds the loop.
  • Load shedding at the overloaded service. Rejecting above measured capacity is what lets it recover; accepting and failing slowly is what prevents recovery.
  • Jittered exponential backoff everywhere — necessary, not sufficient.
  • Make retries observable. Attempt count as a metric distinct from request count, with an alert on the ratio. Without it, dashboards show a traffic spike and nobody knows the traffic is self-inflicted.

Industry example

AWS's Builders' Library discusses timeouts, retries and backoff in these terms, and the retry-budget and token-bucket approaches it describes exist specifically to bound the multiplier rather than merely to space attempts out. The pattern recurs wherever a deep synchronous call graph meets independent per-layer retry policies — most microservice estates — and it is usually discovered the first time a brief dependency blip produces a forty-minute outage.

Failure scenarios

  • The self-sustaining outage above, where the trigger is gone and the overload is not.
  • Retry on a non-idempotent write, duplicating a charge or an order while the caller sees only a timeout.
  • Thundering herd on recovery. The dependency comes back and every client retries in the same second, knocking it down again. Jitter and gradual admission are the controls.
  • Queue-depth retries. A consumer retries a failing message indefinitely, blocking the partition behind it. A dead-letter queue with a bounded attempt count is the fix.
  • Client-side retries you do not control. Mobile apps and browsers retry on their own policy, so a server-side budget cannot bound the total; the only levers are explicit backoff guidance and a fast, cheap rejection.

Trade-offs

Choose Gains Pays
Retry at one layer only Multiplier collapses from product to single factor Transient faults deeper in the graph surface as user errors
Retry budget as a traffic fraction Amplification bounded by construction, globally Some retriable failures are not retried under load
Retry at every layer Maximum transient-fault masking Multiplicative amplification and metastable failure

When not to use it

With a shallow call graph — one or two hops — the multiplier is small and retries are close to free. Three attempts against one dependency is three attempts, and the engineering above is unnecessary ceremony. The concern becomes real at roughly three or more synchronous layers, which is also where it becomes invisible to any single team's review.

And do not remove retries to avoid amplification. Transient faults are real and frequent — a brief network drop, a leader election, an instance recycling — and failing a user request on the first blip is a worse product. The answer is never "no retries"; it is retries at one layer, with a budget, inside a propagated deadline. A team that responds to a cascade by disabling retries everywhere will trade a rare long outage for a constant low-level error rate, which is usually the worse deal.

For asynchronous, queue-based work the calculus also differs: retries there consume throughput rather than holding callers, so a bounded attempt count with a dead-letter queue handles it, and the metastable dynamic is far weaker because nothing is waiting synchronously.

Interview question

Q: A 15-second database slowdown produced a 40-minute outage that did not recover until traffic was shed. Explain what happened and what you change.

What a strong answer covers: computing the multiplier through the layers and naming it as the load source · identifying the state as self-sustaining, so waiting cannot work · distinguishing backoff (spreads attempts in time) from a budget (reduces their number) · retry at one layer, a budget as a traffic fraction, and deadline propagation so no attempt outlives its caller · load shedding as what permits recovery · slow-call breakers, since a slow dependency produces no errors · attempt count as its own metric · and refusing the over-correction of removing retries.

Quick check

Quiz: Four layers each retrying 3 times — how many attempts per user request, and does backoff reduce it? — 81, and no: backoff spaces attempts in time without reducing their number. A retry budget does.

Flashcard: Why does a retry-amplified overload not resolve when the original fault clears? — The retries themselves generate the load that causes the failures that generate the retries. It is a metastable state, and it needs an external intervention — shedding or draining — to break the loop.