advanced 3 min answer

A checkout service calls the platform's configuration service once per request rather than reading at start-up. The config service degrades for 4 minutes with p99 at 12 s and about 30% of calls erroring. Checkout runs at 500 requests per second with a connection pool of 200. Walk through what happens, and what the platform should ship so this coupling cannot recur.

platform slosfail-staticlittles lawcouplingclient library
Show the full answer Hide the answer

Second by second

t+0 to t+1s. Latency rises. By Little's law the in-flight count a caller must hold is arrival rate times latency: 500 rps × 12 s = 6,000 concurrent calls needed, against a pool of 200. The pool is exhausted in well under a second.

t+1s to t+10s. Checkout threads block waiting for a connection. Requests that have nothing to do with configuration now queue behind it, because the exhausted resource is shared across the whole service. Checkout's own latency goes to its timeout, and its error rate approaches 100% - far worse than the 30% error rate of the dependency, which is the detail people get wrong.

t+10s onwards. Retries. Every erroring call is retried by a client that assumes a transient fault, so load on a struggling config service rises by a factor of two to three at the moment it most needs relief. The brownout now has a feedback loop keeping it alive.

Where it amplifies

Checkout is not the only caller. Any service built the same way fails the same way at the same moment, so a 4-minute degradation of one platform component reads on the status page as a correlated outage across unrelated products. The platform's component dashboard meanwhile shows a 4-minute dip against a monthly objective, which is a rounding error. This is why platform availability has to be measured as the consumer experienced it: the platform spent 4 minutes of its own budget and far more of its consumers'.

What the user sees

Checkout failures for the full window plus a recovery tail, because the pool has to drain and the retry backlog has to clear. Expect user-visible impact of roughly two to three times the duration of the underlying event.

What stops it

A dependency only enters your availability chain if you call it synchronously in the request path. The fix is to remove it from that path, and the platform should ship the mechanism rather than the advice:

  1. A client library that reads configuration at start-up, caches it in memory, refreshes in the background, and serves the last known good value when the refresh fails. Fail-static, with a bounded staleness the platform states - say 5 minutes - and a metric for cache age.
  2. A start-up behaviour that is explicit: a missing config at boot either blocks the start (safe, but makes a config outage a scaling outage) or loads a baked-in default. Choose per value and write it down.
  3. Consumption terms published with the SLO. The objective should read "99.9% availability, read at start-up with background refresh, cache age tolerated to 5 minutes". A number without a stated consumption pattern is not a usable commitment.
  4. A game day that blackholes the config service and confirms no consumer's error rate moves - the only evidence that the library is used as intended - plus the share of callers on the fail-static version published as a platform metric, since that number is the platform's real exposure.

When this is the wrong answer

Fail-static is wrong where staleness has a safety or correctness cost: a kill switch, an entitlement check, a rate limit that exists to stop abuse. Those belong in the request path. Keep them there with a hard timeout well under the caller's remaining budget - tens of milliseconds - a local default for the timeout case, and a circuit breaker that trips on slow calls rather than only on errors.

Common weak answers

  • "Add a retry with backoff." The dependency is slow, not down. Retries add load to the thing that is struggling and do not refill an exhausted pool.
  • "Raise the connection pool to 6,000." You have now sized for the failure mode and moved the exhaustion to memory and to the config service itself.
  • "Put the config service behind a cache in the platform." A shared cache in front of a shared service keeps the synchronous hop and adds a component. The caching that fixes this is in the caller's process.