An order service calls a pricing service. To protect itself it has three retries with jitter, a breaker that opens at a 50% error rate over 10 seconds, a 30-second-TTL fallback cache, and a bulkhead of 20 threads for pricing calls. At 09:00 pricing starts answering in 4 seconds instead of 80 ms, with no errors at all. Pricing normally serves 250 requests per second. Walk through what each of the four mechanisms does, and say which one makes the outage worse.
Show the full answer Hide the answer
Second by second
The bulkhead saturates first, and it saturates immediately. Concurrency needed equals arrival rate times latency: at 250 rps and 80 ms the service needed 20 concurrent slots, which is exactly what the bulkhead was sized for. At 4 seconds it needs 250 x 4 = 1,000 slots. The bulkhead holds 20, so it admits 20 / 4 = 5 pricing calls per second, which is 2% of demand. Within the first second, 98% of pricing calls are being rejected at the bulkhead or queued behind it.
The breaker never opens. Slow is not an error. A 4-second success is a success, so the error rate stays near zero and the 50%-over-10-seconds condition is never met. The mechanism everyone assumes saved them contributes nothing to this incident, which is the single most common gap in a resilience stack.
The retries make it worse. Every bulkhead rejection and every timeout is retried three times, so offered load on a dependency that is already saturated triples. This is the mechanism that converts a slowdown into an outage: the dependency's recovery requires offered load to fall, and the retry policy guarantees it rises.
The fallback cache does almost nothing. A 30-second TTL covers only keys requested in the last 30 seconds, and it was populated by the same calls that are now failing. Hit rate collapses precisely when it is needed. A fallback populated by the path that is failing is not a fallback.
What the user sees
Order submission hangs for as long as the caller's own timeout allows, then fails with a generic error. Order service threads are consumed waiting, so endpoints that never touch pricing start timing out too - the bulkhead was supposed to prevent that, and it does for the thread pool it guards, but not for the request threads blocked upstream of it.
The one change that has to come first
A timeout below the caller's remaining budget. If the user-facing budget is one second, pricing gets 300 ms. Nothing else in the stack can act until slow becomes an error, because every other mechanism keys off failure. With the timeout in place: the breaker sees errors and opens, the retry policy has something to decide about, and the bulkhead stops holding threads for four seconds each.
Then, in order: trip the breaker on slow-call rate rather than error rate, which Resilience4j supports directly and which catches exactly this incident; cap retries with a retry budget - a common setting is retries limited to about 10% of total requests - so a saturated dependency cannot be flooded; and populate the fallback out of band with a TTL measured in hours, so it holds data when the live path does not.
When not to carry all four
Four interacting mechanisms are past the point where a human can predict their joint behaviour, which is what this incident demonstrates. For a dependency that is either up or hard-down - a connection refused, a DNS failure - a timeout plus an error-rate breaker is the complete answer, and the bulkhead and the fallback cache are complexity with no payoff. Add each mechanism only against a named failure you have actually seen, and write down what it is expected to do in response. A mechanism nobody can describe the behaviour of is a liability on the next incident, not an asset.