advanced 3 min answer

A ticketing platform at Ticketmaster's on-sale peaks routes everything through one gateway. Gateway to BFF to inventory is three hops and each hop retries up to three times on a timeout or a 5xx. Inventory's p99 crosses its 400 ms timeout at 09:00:02. Walk the next ninety seconds and say where the limit actually has to live.

api gatewayretriesretry budgetadmission controlamplification
Show the full answer Hide the answer

Second by second

  • 09:00:00 - on-sale opens. 180,000 requests a second arrive at the gateway. Everything is fine.
  • 09:00:02 - inventory's p99 crosses 400 ms, so about 1% of calls time out. Each hop retries, so the leaf sees roughly 1.03x its normal load. Nothing alerts. The amplification is present and latent.
  • 09:00:06 - inventory's queue grows, its p50 crosses 400 ms, and now every call times out. Three attempts at each of three hops is 27x offered load: 4.8 million requests a second at a service provisioned for 200,000. That multiplier is the well-known part of this failure.
  • 09:00:09 - the part that is usually missed. A retry occupies a gateway worker for the full timeout, so the gateway's own in-flight count triples. With 4,000 workers and a 400 ms timeout the gateway can hold about 10,000 concurrent calls; it is now full of retries for one route and starts returning 503 for the seat map and the account service, which have nothing wrong with them. The failure has jumped routes.
  • 09:00:30 - retries no longer come only from servers. The mobile app has its own retry loop and users are pressing refresh. Client-generated load keeps climbing after every server-side retry has been suppressed, and no server-side control bounds it.
  • 09:01:30 - nothing has recovered. Inventory cannot drain a backlog that is still arriving at many times its capacity, and each request it finally answers is for a caller who left a minute ago.

Where the limit has to live

At the hop adjacent to the failing dependency, expressed as a fraction of that dependency's own traffic, and held per upstream cluster. Three properties, each load-bearing.

A budget rather than a count, because a budget composes. With a per-hop budget of fraction b over h hops, worst-case amplification is (1+b)^h: at 20% across three hops that is 1.73x, which inventory absorbs. Envoy's retry budget works this way and defaults to 20% of active plus pending requests (2026 documentation). A per-request count of three gives 27x and no amount of tuning the count changes the shape.

Per upstream cluster, because a single gateway-wide budget lets the failing route consume the allowance that healthy routes needed.

Not only at the gateway, because the gateway's budget bounds only the gateway's retries and is blind to what the BFF does. The hop nearest the failure is also the only one with evidence about that dependency's health.

Two things sit alongside it. Retries must be charged against inbound concurrency, or the gateway saturates before the leaf does, which is the 09:00:09 step above. And the client is bounded only by a signal it is told to honour: 429 or 503 with Retry-After, plus an app that actually respects it, which is a client release and therefore a decision to make before the on-sale rather than during it.

What would have to be true for it to self-heal

Offered load has to fall below capacity without anyone intervening. Three mechanisms together do that: the per-cluster retry budget, deadline propagation so that work whose caller has already gone is dropped rather than started, and per-route shedding at the gateway so inventory's collapse cannot consume the gateway. Add capacity and you get the same curve at a higher number.

When not to add a budget

When the dependency fails fast rather than slowly, and the chain is one hop deep. A connection refused in under a millisecond, a DNS miss, a 503 from a healthy load balancer with no backend: three attempts with full jitter cost almost nothing, because a failed attempt does not occupy a worker for the timeout. Retry amplification is a latency-shaped failure. Budgets are what you need when the dependency stays up and gets slow, which is the common case and the one that takes the platform down.