A checkout sequence diagram has six participants and every arrow is a plain synchronous call with no timeout written on it. In production the fourth participant starts answering in 9 seconds instead of 200 ms and never returns an error. What happens second by second, what stops it, and what should the diagram have carried?
Show the full answer Hide the answer
Second by second
Take a front end with a 200-thread request pool serving 40 requests per second, holding a thread for 300 ms in the normal case. By Little's law that is 40 × 0.3 = 12 threads busy out of 200, which is why nobody has ever thought about the pool.
The dependency slows to 9 s. Demand for concurrency becomes 40 × 9 = 360 threads against a pool of 200.
- 0 to 5 s: threads accumulate. Latency rises for every endpoint, including the ones that never touch the slow participant, because the pool is shared.
- ~5 s: the pool is exhausted. New requests queue at the socket. p99 for the whole service is now the dependency's latency plus queue wait.
- ~10 s: the health check shares the pool, so it cannot be answered. The load balancer marks the instance unhealthy and removes it.
- Next interval: the removed instance's share of traffic lands on its siblings, which are already at the same arithmetic. They exhaust in turn. The service is fully down because of a dependency that returned zero errors.
Where it amplifies
Two multipliers. Retries: a client-side retry on a slow call doubles or triples offered concurrency exactly when the pool is short. Health-check coupling: a check that competes for the same threads converts degradation into removal, which converts removal into cascade.
What the user sees
Not a payment error. A checkout that spins and a browser tab that eventually times out, which is why support tickets say "the site is slow" while the dependency's own dashboard is green. The dependency reports success on every call it answers.
What stops it
- A deadline set at the edge and propagated — say 2 s for checkout — so every hop knows its remaining budget and abandons rather than waiting. This is the mechanism; the rest are refinements.
- A per-dependency concurrency limit: allow the slow participant at most 25 of the 200 threads. The rest of the service survives with one broken feature.
- A breaker that trips on slow-call rate, not error rate. An error-rate breaker never opens here, because there are no errors.
- A health check on its own thread so slow does not become removed.
Self-healing requires bounded concurrency plus a retry budget. Without both, the system has no mechanism to reduce its own offered load and will not recover until the dependency does.
What the diagram must carry
Every synchronous arrow gets a timeout, and the sum along the critical path is compared against the edge deadline. If the written timeouts add to more than the deadline, the diagram is already describing an outage and the review should stop there. Mark which participants are on the user's critical path and which could be made asynchronous — in checkout, the receipt email and the loyalty update usually can be, and moving them off the path removes two of the six arrows from the budget.
When this is the wrong answer
For a flow with two participants and a hard dependency, a timeout alone is enough and the bulkhead is ceremony. Concurrency limits earn their keep from about three dependencies, where one being slow should not spend the whole pool.