Service Mesh Platform  ·  View 20 of 31  ·  5 · Runtime

A Slow Dependency, Contained

A dependency's latency rises tenfold. Four mechanisms stop three retry attempts per call from tripling load on the one service already failing.

Editable source SVG draw.io All views
Dependency slows Retries begin Budget reached Shedding Recovery Caller proxies p99 280 ms → 2 s Retry 503 · GET only Budget 20% spent Fast 503 to app Budget refills Callee proxies Pending grows Outlier ejects 2 pods Concurrency at limit Overflow shed Ejections expire Deadline Caller allows 800 ms 250 ms per try No retry past deadline Signals Latency alert Retry ratio 0.2 retry_overflow Origin: mesh limit Budget refills A Slow Dependency, and Why It Does Not Become an Outage Without the budget, three attempts per call triples load on the one service already failing. With it, load rises by at most a fifth. v 1.0 · owner Platform Networking Architecture · date 2026-09

Decisions

  • Retries are capped by a budget of 20% of active requests per destination, not only by an attempt count. Attempt counts bound one call; budgets bound the fleet.
  • No retry starts if the caller's remaining deadline cannot cover a per-try timeout. The deadline travels as a header, and the inbound proxy caps its own route timeout by it.
  • The callee's proxy sheds at its concurrency limit with an immediate 503 marked as a mesh limit. The caller's app sees fast failure; operators see which layer said no.

To prove

  • Istio's API does not expose Envoy's retry budget or its expected-timeout handling on every release. Both are applied through one generated, version-pinned EnvoyFilter covered by the upgrade conformance suite until the API catches up.

Assumptions

  • Retryable conditions are connection failure, reset and 503 on GET, HEAD and declared idempotent methods; three attempts; 25 ms base backoff with full jitter.