Distributed Systems 09 Oct 2026 30 min read

Choosing the backend

How production systems decide which of N supposedly interchangeable replicas gets each request, and why the selection algorithm keeps manufacturing the hot spots and black holes it exists to prevent.

Reconstructs the backend-selection stack from Google's SRE book and Prequal paper, Netflix's edge redesign, Twitter's aperture work, Uber's dynamic subsetting, gRPC's design records (including one subsetting proposal closed unmerged), Envoy and Finagle's own documentation, and incidents at Slack, GitHub and Heroku. After reading, an architect can choose a selection policy and its guards deliberately, write the trust contract a self-reported load signal needs, and recognise the four feedback loops through which a chooser creates imbalance: rewarding failure, herding on stale signals, serving a stale pool view, and the rebalance itself.

The finding that surprised me

The mature implementations agree the balancer's job is not to find the best backend but to bound how much it trusts every signal that claims to know one: least-loaded routing without an error term sends the most traffic to a server failing fastest, because errors are cheaper to produce than answers.

What you get out of it

  • Every busy-ness score needs an error term: a failing server completes 'work' faster than a healthy one, so least-loaded policies route traffic into it (SRE book mechanism; defended against in gRFC A58, Netflix guardrails, Envoy outlier ejection; no public dated postmortem names it, which is itself a gap worth knowing).
  • The per-request pick must contain private randomness; shared or stale load signals may only shape weights and subsets slowly, or the choosers herd (Brooker, C3, Envoy's stated reason for P2C).
  • Two sampled choices capture the win: exponential improvement over random, while a third choice adds only a constant factor (Mitzenmacher 2001), which is why P2C is the convergent default in Envoy and Finagle.
  • Subsetting is a three-way trade measured as connections saved, worst-replica load, and churn under rollout: Twitter, Google and gRPC each weighed it differently and landed on different algorithms.
  • Membership change is the dangerous moment: a balancer reload synchronises reconnects (GitHub 2026), a cold host gets a full share without slow start (Envoy), and slow start itself admits it cannot help when the whole fleet is new.

Scope

Why this, now. The field is re-litigating its signal right now: Google published Prequal in 2024 arguing CPU load is the wrong thing to balance, gRPC merged its weighted-round-robin trust contract in 2023 and settled its subsetting argument only in 2024-2026, and GitHub published two load-balancer incidents in the last fourteen months.

What it does not cover. Consistent hashing for cache affinity, L4 and anycast packet steering (Maglev, Unimog), cross-region traffic steering, overload control after the request lands, and failure detection itself, which this site's 'deciding a server is dead' guide covers.

Open the field guide → Self-contained: it loads nothing at read time, follows your system theme, and prints cleanly.