advanced 3 min answer

At 10:04 one dependency slows from 40 ms to 900 ms. By 10:06 the service is saturated, the dependency is receiving 14 times its normal request rate, and nothing in the code review that introduced the retries looks wrong. What failed, and which design decision made it possible?

retriesdecoratoramplificationcascading-failureresilience
Show the full answer Hide the answer

The trigger

A dependency's latency rose. Nothing failed outright. That is the hard case, because the mechanisms built to protect the system key on errors, and a slow dependency produces none until a timeout converts it into one.

Why it propagated

Three independent layers each add retries, and each was reviewed on its own merits.

  • The HTTP client library is configured with 3 attempts, a sensible default someone set in a shared module two years ago.
  • A RetryingRepository decorator wraps the client, also with 3 attempts, added when a flaky integration was causing support tickets.
  • The use case is itself retried by the message consumer, 3 attempts, because at-least-once delivery requires it.

Multiplication, not addition: 3 × 3 × 3 = 27 attempts for one logical request. At normal latency this is invisible, because the first attempt almost always succeeds. Under slowness, every attempt reaches its timeout and the next layer begins again, so the dependency's load rises by roughly an order of magnitude at the exact moment it is least able to absorb it. The observed 14× is the partial version of this, with some requests succeeding on a later attempt.

The arithmetic is not new: Google's SRE book (2016) describes the same amplification and prescribes a retry budget as the control. The second multiplier is threads, and it costs you the whole service: each in-flight attempt occupies a worker for the duration of its timeout. If the outer timeout is longer than the sum of the inner ones, which it usually is because each was chosen locally, the service exhausts its worker pool and stops serving traffic that has nothing to do with the slow dependency.

Why detection lagged

Dashboards showed the dependency's error rate near zero and its latency elevated. They did not show the ratio of outbound attempts to inbound requests, which is the signal that would have named this in seconds. Each layer's own retry metric looked reasonable; only the product is pathological, and nobody owns the product.

The structural fix versus the tempting local fix

The tempting fix is to reduce one layer's retries from 3 to 2, which halves nothing and leaves the pattern intact.

The structural fix:

  1. Retry at exactly one layer, and make the others fail fast. The right layer is the one that knows the request's business meaning and its deadline, which is usually the outermost.
  2. Give the request a retry budget rather than a per-call count. A budget expressed as a percentage of request volume (for example, retries capped at 10% of requests) makes amplification impossible by construction, because the cap is global rather than per call site.
  3. Propagate a deadline. Each layer works with the remaining budget and refuses work it cannot finish, so an inner retry cannot outlive the caller's patience.
  4. Trip on slow calls, not only on errors. A breaker keyed on error rate never opens here; one keyed on the share of calls exceeding a latency threshold does.
  5. Alert on attempts per inbound request, per dependency. A ratio above about 1.1 in steady state is a design problem waiting for a bad afternoon.

The general lesson

Composable patterns compose their failure modes, and a decorator stack hides the composition. Each layer is locally correct, individually reviewable, and invisible to the others. Any cross-cutting behaviour that can be stacked — retries, caches, timeouts, circuit breakers, rate limiters — needs one place where the total is declared and tested. Where the pattern must stay distributed, write the arithmetic into the code review checklist: a reviewer approving a retry has to say what the total attempt count for the path becomes.

When not to collapse to a single retry layer

Retrying at several layers is defensible when the layers protect genuinely different failures and the inner one is bounded and fast: a single immediate reconnect inside a connection pool, under a 50 ms cap, beneath a business-level retry with a 2-second budget. The test is whether the worst-case total is bounded and written down. If nobody in the room can state the maximum number of attempts a request can cause, the stack is unsafe regardless of how reasonable each layer looks.