Choosing the backend  / field guide
Practitioner field guide · 2026-10-09

Choosing the backend: why the picker keeps creating the hot spot it exists to prevent

A service runs N supposedly interchangeable replicas, and every request must land on exactly one of them, chosen with information that is stale, local, or both. This guide reconstructs how that choice is actually made at Google, Netflix, Twitter, Uber, Slack, GitHub and Heroku, from their own postmortems, design records and rejected proposals. Afterwards you can pick a selection policy and its guards deliberately, and recognise the four ways the chooser manufactures the imbalance it was installed to remove.

30 primary sources 14 organisations 6 published incidents Evidence through October 2026 Read: 25 min
01

The territory

One request, many equal replicas, and a chooser working from signals that lie. Who has solved this in production, and what it costs when the chooser is wrong.

>99%
Connection count cut by subsetting at Twitter, before it wrecked request fairness
1M+ rps
Netflix edge traffic when it replaced round robin with a two-signal choice-of-2
1.8%
Of GitHub Actions runs delayed, 15 min average, by a single connection rebalancing event
10 s
Blackout before gRPC trusts a new backend's self-reported load at all

State the problem without naming a product: many copies of the same server, one incoming request, and a component that must assign the request to exactly one copy using whatever it can observe. The copies are never actually equal. Netflix's edge team listed the reasons in 2018: servers are cold right after deployments and autoscaling, they stall during configuration updates and large garbage-collection pauses, and some hardware simply runs permanently slower than the rest. At their scale, over a million requests per second, even a low rate of load-related errors was worth a redesign.

The field looks settled from a distance: round robin, least connections, done. Up close it is one of the most actively contested corners of production engineering. AWS did not offer anything beyond round robin on its Application Load Balancer until November 2019. Envoy shipped latency-aware selection only as a build-time contrib extension. gRPC merged a weighted round robin design in 2023, then argued through two subsetting proposals, closing the first unmerged. And in 2024 Google published a paper about the balancer in front of YouTube whose title is a flat rejection of received wisdom: "Load is not what you should balance." Teams are still changing their answer, which means the decisions are still live.

The single idea that organises everything in this guide: the chooser is a feedback controller, not a dispatcher. Every signal it consumes, a connection count, a latency average, a server's self-reported CPU, is also a path through which its own past decisions, other balancers' decisions, and the failure modes of the fleet flow back into its next decision. Each of the four failure classes in section 4 is one of those feedback paths closing into a loop. The craft is not picking the cleverest scoring function; it is bounding how much the controller is allowed to trust each signal, and for how long.

Figure 1 · The chooser is a control loop

in-flight, latency,
self-reports, probes

health verdicts,
registrations

Incoming
request

Pool view

Subset

Pick one:
random / RR /
P2C / WRR

Chosen
replica

in-flight, latency,
self-reports, probes

health verdicts,
registrations

Incoming
request

Pool view

Subset

Pick one:
random / RR /
P2C / WRR

Chosen
replica

Every arrow on the right flows back into the next pick, which is why a selection policy can amplify the very imbalance it measures. Reconstructed from Netflix, 2018 and Google's Prequal paper, 2024.
Diagram source

Scope. This guide covers the selection decision: which live replica gets this request, made by an L7 proxy, a client library or a sidecar. It deliberately excludes consistent hashing for cache affinity, L4 and anycast packet steering (Maglev, Unimog), global traffic steering across regions, overload control once the request has landed, and the separate question of deciding a server is dead, which this site's failure-detection guide already covers. Health checking appears here only where the verdict feeds the pick.

02

How it is actually built

The same four-stage shape appears in Google's GSLB, Netflix's Zuul, Twitter's Finagle, Uber's mesh, Envoy and gRPC. The divergence is in which signal each stage trusts.

Figure 2 · Reference architecture of the selection stack

Backend fleet

Balancer instance

Control plane

weights

in-flight, latency, utilization
header, probe answers

Service discovery
and registration

Load-report aggregation
(Uber control plane, xDS LRS)

Pool view

Subsetter

Selector
(P2C on in-flight is
the common default)

Guards: slow start,
outlier ejection,
panic threshold

Replica

Replica
(new: warming)

Backend fleet

Balancer instance

Control plane

weights

in-flight, latency, utilization
header, probe answers

Service discovery
and registration

Load-report aggregation
(Uber control plane, xDS LRS)

Pool view

Subsetter

Selector
(P2C on in-flight is
the common default)

Guards: slow start,
outlier ejection,
panic threshold

Replica

Replica
(new: warming)

The pick itself is always local and fast; everything shared or slow (discovery, load aggregation) feeds it asynchronously. Reconstructed from Google's SRE book, Uber, 2022 and the Envoy architecture docs.
Diagram source

Four stages recur in every published system. First a pool view: what the balancer currently believes the fleet is, assembled from discovery and health verdicts. It is a belief, not a fact, and section 4 shows what happens when it drifts. Second a subsetter, present once fleets get large: each client restricts itself to a window of the pool so that connection counts stay affordable. Third the selector proper, which scores candidates and picks one. Fourth a layer of guards that override the score in specific situations: a new host gets a ramp, a suspicious host gets ejected, and a mostly-unhealthy pool triggers a deliberate change of policy.

The selector: sample two, compare, pick

The convergent default is power of two choices over a local in-flight counter: pick two random candidates, send to the one with fewer outstanding requests. Finagle's documentation calls it the default balancer for all clients; Envoy's least-request policy is "an O(1) algorithm which selects N random available hosts (2 by default)" and cites Mitzenmacher directly, choosing P2C for its "resistance to herding behavior". Netflix's edge rebuilt the same shape with a richer score.

Runs this way at: Twitter (Finagle), Envoy, Netflix

The signals: three families, three trust problems

Local counters (in-flight requests) are always fresh but see only this balancer's traffic. Latency averages (peak EWMA) react fast to slow hosts but, per Finagle's own docs, break under long-polling clients. Server-reported load (utilization headers, ORCA reports, probe responses) sees the whole truth but arrives late and can lie, which is why gRFC A58 wraps it in a 10-second blackout, a 3-minute expiry and an error penalty before the scheduler may act on it.

Signal contracts: gRFC A58, Prequal, Finagle

The guards: overrides, not scores

Envoy's slow start scales a new endpoint's weight up over a configured window, because otherwise a new host immediately receives a full proportional share. Outlier ejection removes hosts that keep erroring, covering the blind spot of a selector that finds errors attractive. The panic threshold inverts the whole policy: once fewer than 50% of hosts look available, Envoy stops believing health status entirely rather than concentrating all traffic on the survivors. Finagle ships the same idea as panic mode.

Sources: Envoy slow start, Envoy panic threshold

Two structural divergences matter more than any scoring detail. Where the chooser lives: a dedicated proxy tier (Heroku's router, GitHub's internal balancing layer, ALB), a client library (Finagle, gRPC), or an on-host sidecar (Uber's mesh, Envoy). Client-side picking removes a hop and a point of failure but multiplies the number of independent, non-cooperating controllers; Twitter's SREcon talk is explicit that Finagle's balancing is "non-cooperative, client-side", and the whole aperture design exists to make thousands of selfish choosers produce a fair global outcome. Whether request cost is visible: balancing connections is meaningless once a protocol multiplexes requests over one long-lived connection. The Kubernetes blog's gRPC demo makes this concrete: HTTP/2 keeps a single TCP connection per client, so connection-level balancing parked every request on one pod while the rest idled.

One more recurring component deserves its name said out loud: the load-report aggregator. Uber's control plane collects per-task load and recomputes subset sizes in close to real time; xDS has a load reporting service; Prequal keeps a pool of probe responses per client. Across all three, the aggregated signal feeds weights and subset shapes, never the per-request pick itself, which stays on local state. Nobody published a system where a shared load database sits in the per-request path. That absence is the architecture speaking: a shared signal at pick time is a synchronised signal, and synchronised choosers herd.

03

The decisions that matter

Four forks, each with a published team on both sides, and the condition that moves the answer.

Decision: which signal is the pick allowed to trust?

Chosen
  • Requests-in-flight plus latency, locally observed. Google's Prequal selects on "estimated latency and active requests-in-flight" and explicitly not CPU; Netflix picks primarily on the balancer's own view of a server, secondarily on the server's self-report; Finagle defaults to least-loaded over P2C.
Rejected
  • CPU-based weights as the sole signal. The Prequal paper's title is the rejection: queued work and latency predict trouble; CPU utilization trails it.
  • Cached global load snapshots. Brooker, twice (2012 and 2024): "the stale data quickly causes bad decisions to be made."
Flips when
  • Request costs vary wildly and backends report calibrated load under a contract with teeth: gRFC A58's weighted round robin is exactly this, and it is viable because of its blackout, expiry and error-penalty guards, not despite them.

Decision: sample two candidates, or rank the whole pool?

Chosen
  • Power of two choices. Mitzenmacher's result is the reason it is enough: two choices capture an exponential improvement over one, and a third adds only a constant factor. Envoy made it the least-request default; Brooker reports deploying best-of-k "many times in large-scale production systems".
Rejected
  • A globally ordered structure. Finagle's own docs explain why its older heap balancer lost the default: the heap "must be updated atomically by each request and thus represents a highly contended resource", and swapping load metrics is hard.
  • Full O(N) scans: Envoy's docs note P2C is "nearly as good as an O(N) full scan".
Flips when
  • Concurrency is too low for the load metric to mean anything: Finagle warns that "without sufficient concurrent load, the previous distributors can degrade to random selection", which is what its aperture band targets fix.
  • Bins are very small, Brooker's one named exception for best-of-k.

Decision: subset the pool, and by what rule?

Chosen
  • Deterministic windows. Twitter's deterministic aperture places clients and backends on two rings at equal intervals and runs P2C inside the window; Google's SRE book published a deterministic round-based subsetter years earlier.
  • Dynamic sizing. Uber recomputes subset size per service from real-time load reports, because a fixed size that fits one caller starves or swamps another.
Rejected
  • Full mesh: Twitter's random aperture existed because connecting everyone to everyone was already unaffordable; the subsetting cut aggregate connections by over 99%.
  • Plain random subsets: the same change made "a mess of our request distribution", with visible bands of load and hot instances near 400 rps.
Flips when
  • Connection churn during rollouts dominates: Google's 2023 CACM account describes replacing deterministic subsetting with an algorithm designed to reduce churn, and gRPC closed its deterministic subsetting proposal unmerged, then shipped random subsetting with rendezvous hashing instead. Determinism buys fairness and costs churn; which one hurts more decides.

Figure 3 · A decision path through the published choices

yes: HTTP/2, gRPC

no: HTTP/1.1 pools

yes

no

yes, with blackout,
expiry, error penalty

no

Protocol multiplexes requests
on one connection?

Balance requests, not connections:
L7 proxy or client-side balancer

Hundreds of backends
per client?

Subset first: deterministic window, or
random + rendezvous if churn beats skew

Trustworthy reported
load signal?

WRR on reported load (gRFC A58)
or probe pool (Prequal)

P2C on local in-flight,
slow start, outlier ejection

yes: HTTP/2, gRPC

no: HTTP/1.1 pools

yes

no

yes, with blackout,
expiry, error penalty

no

Protocol multiplexes requests
on one connection?

Balance requests, not connections:
L7 proxy or client-side balancer

Hundreds of backends
per client?

Subset first: deterministic window, or
random + rendezvous if churn beats skew

Trustworthy reported
load signal?

WRR on reported load (gRFC A58)
or probe pool (Prequal)

P2C on local in-flight,
slow start, outlier ejection

Terminal nodes are configurations real teams run, not "it depends". The first question outranks the algorithm choice entirely, per the Kubernetes gRPC post.
Diagram source

Heroku's 2013 crisis is the cautionary tale for the first row of the table below, and it is often retold wrong. Random routing was not a bug; Heroku's FAQ defends it on availability and stateless scaling grounds, saying the team found nothing that "beat the simplicity and robustness of random routing to back-ends that support multiple concurrent connections". The failure was the qualifier: Rails on Thin served one request at a time, so random assignment built queues behind busy dynos while others idled, and the pain concentrated exactly on "Rails apps running on Thin, with six or more dynos, serving 1k requests per minute or more". The press story was marketing ("intelligent routing" stayed on the website after the mid-2010 switch, per The Register); the engineering lesson is that random is a legitimate policy whose precondition is concurrent backends, and the precondition is the part teams forget to check.

DecisionChosenRejectedBecauseEvidence
Queue awareness at all?Random routing (Heroku)Global per-app request queueAvailability and stateless scaling of the router tierHeroku FAQ, 2013
Default ALB algorithmRound robin until 2019, then LOR optionLOR as defaultAWS added LOR for varied request costs and target churn; RR remains defaultAWS, 2019
Scan or sampleP2C (Envoy, Finagle default)Heap / full scanContention and herding resistance; two choices capture the winEnvoy docs, Mitzenmacher, 2001
Selection signalIn-flight + latency (Prequal, Netflix)CPU as sole signalQueued work predicts trouble before CPU doesNSDI '24
Self-reported weightsWRR with blackout, expiry, error penaltyTrusting reports immediately and foreverNew and stale reports are the two ways the signal liesgRFC A58, 2023
Subsetting rule in gRPCRandom + rendezvous hashingDeterministic subsetting (A68 draft)First proposal closed unmerged; churn and complexity concerns wonPR #383, PR #423
New-host treatmentSlow start window (Envoy), warm-up before trust (A58 blackout)Full share at joinCold hosts time out and lose data at full shareEnvoy docs
Mostly-unhealthy poolPanic: ignore health or fail fastConcentrating all traffic on the last healthy hostsPrevents the survivors being crushed in cascadeEnvoy docs
04

What broke in production

Four failure classes cover every published incident this hunt found. Each is one of the control loop's feedback paths closing.

Class 1: the chooser rewards failure. A server that fails requests is faster than a server that serves them, so every load metric built on "busy-ness" scores the broken server as the best choice. Google's SRE book states it plainly for least-loaded round robin: a seriously unhealthy task may serve 100% errors, those errors "may have very low latency" because "it's frequently significantly faster to just return an 'I'm unhealthy!' error than to actually process a request", and clients then send the unhealthy task a very large amount of traffic. This is the one class with no named public postmortem in this corpus, and that gap is itself a finding: the mechanism is documented in a book, defended against in gRFC A58's error penalty, Netflix's guardrails and Envoy's outlier ejection, but no organisation has published a dated incident report saying it bit them. Either the defences work, or the incidents get filed under their symptoms. Design as if it will happen: any score derived from completion speed needs an error term before it is safe.

Figure 4 · The error sinkhole under least-request

Replica B: failing, 2ms errorsReplica A: healthy,120 msBalancer,least-requestReplica B: failing, 2ms errorsReplica A: healthy,120 msBalancer,least-requestB looks idle againErrors are cheaper than work, so B wins nearly every two-waycomparisonrequest 1 (A in-flight: 1)request 2 (B in-flight: 1)500 after 2 ms (B in-flight: 0)request 3500 after 2 msrequest 4200 after 120 ms
Replica B: failing, 2ms errorsReplica A: healthy,120 msBalancer,least-requestReplica B: failing, 2ms errorsReplica A: healthy,120 msBalancer,least-requestB looks idle againErrors are cheaper than work, so B wins nearly every two-waycomparisonrequest 1 (A in-flight: 1)request 2 (B in-flight: 1)500 after 2 ms (B in-flight: 0)request 3500 after 2 msrequest 4200 after 120 ms
The failing replica finishes "work" sixty times faster than the healthy one, so its in-flight count is almost always lower and it wins almost every comparison. Mechanism from Google's SRE book, ch. 20.
Diagram source

Class 2: herding on shared or stale signals. When many choosers act on the same delayed information, they all draw the same conclusion at once, and the "least loaded" server is buried by the crowd that identified it. Brooker's 2012 and 2024 posts give the mechanism; the C3 paper built its Cassandra replica selector specifically to rank fast servers with growing queues lower, to avoid herd behaviour, and cut p99.9 latency by up to 3x; Envoy's docs justify P2C by its "resistance to herding behavior". The rule the sources converge on: shared signals may shape weights and subsets slowly, but the per-request pick must contain private randomness.

Class 3: the pool view is not reality. The selector can only distribute across the fleet it believes exists. Slack's May 2020 outage is the canonical account: the balancer's belief aged eight hours while autoscaling churned the real fleet underneath it. Class 4: the rebalance is itself an event. Changing the assignment, on a deploy, a scale event or an LB reload, moves load in bulk, synchronously, which is exactly the traffic pattern the steady-state algorithm never sees. GitHub has published three incidents in this class in under two years; Envoy's slow start docs admit the guard does not help "when all the endpoints are relatively new e.g. new deployment in Kubernetes", because a ramp that applies to everyone ramps no one relative to anyone else.

Postmortem

Slack: the balancer served yesterday's fleet

AssumptionHAProxy's view of the webapp fleet tracks the real fleet closely enough to survive a routine scale-down.
What happenedMost HAProxy instances "were stuck with full and stale backend state" more than eight hours old; autoscaling terminates oldest instances first, so the stale slots pointed disproportionately at hosts that no longer existed. The evening scale-down removed the last real capacity the stale view still knew about.
Blast radiusFull user-facing outage 4:45 to 5:33 p.m. PDT on 2020-05-12; degradation began hours earlier. The alert for exactly this condition existed and "wasn't working as intended".
FixRolling restart of the HAProxy fleet to reload state; monitoring on the staleness itself.
Design ruleThe pool view is state with an age. Measure its age, alert on it, and make "rebuild the view from scratch" a routine operation rather than an incident response.
Postmortem

GitHub: one rebalance, 15-minute delays

AssumptionConnection rebalancing inside the LB layer is an invisible housekeeping operation.
What happened"A connection rebalancing event in our internal load balancing layer... temporarily created uneven traffic distribution across sites and led to request throttling." Clients reconnected in bulk and landed unevenly.
Blast radius2026-02-23, 15:00 to 17:00 UTC; 1.8% of Actions workflow runs delayed, average 15 minutes.
FixTune rebalancing "so client reconnections spread out more gradually during load balancer reloads"; site-level traffic affinity; monitoring for capacity imbalance.
Design ruleEvery LB reload is a synchronized reconnect of your most active clients. Jitter the reconnects by design, and treat reload tooling as a traffic event with a blast-radius budget.
Postmortem

GitHub: retries ate the balancer

AssumptionRetry logic in front of flaky LB-to-backend connectivity makes the flakiness harmless.
What happenedIntermittent connectivity between load balancers and search hosts was first masked by retries; the accumulated retry queues then exhausted the load balancers themselves, and the masking layer became the failing layer.
Blast radius2025-08-12, 13:30 to 17:14 UTC; up to 75% of search queries failed at peak; index updates lagged up to 100 minutes.
FixSlowed the indexing pipeline to shed load, tuned the search cluster's load balancer, improved monitors and playbooks.
Design ruleThe balancer holds state per outstanding attempt, so retries multiply its memory of the incident. Budget retries against the balancer's capacity, not only the backend's.
Postmortem

Heroku: random routing meets one-request backends

AssumptionBackends accept multiple concurrent requests, so random assignment spreads load acceptably.
What happenedRails on Thin served one request at a time. Random routing queued requests behind busy dynos while others idled, and the dyno-side queueing was invisible in New Relic's router-reported queue time, producing "unexplained high latencies, mismatched queuing metrics, and differences between documented and observed behavior".
Blast radiusYears of degraded latency for the affected class, surfacing publicly in February 2013; worst for Rails apps on Thin with six or more dynos above 1k requests per minute.
FixHonest documentation and metrics (queue time measured at the dyno), guidance to concurrent servers such as Unicorn; random routing itself was kept, deliberately.
Design ruleEvery selection policy has a precondition written in its proof sketch. Random requires concurrent backends; least-request requires honest completion costs. Verify the precondition, not the policy.
Case study

Twitter: subsetting saved connections, shredded fairness

AssumptionIf every client picks a random subset, the overlaps average out and load stays even.
What happenedRandom aperture cut aggregate connections by over 99%, but random subset membership is lumpy: "we've made a mess of our request distribution", with distinct bands of load and hot instances approaching 400 rps while others sat nearly idle.
Blast radiusNo outage; the cost was capacity planning. Service owners had to provision for the hottest band, not the average.
FixDeterministic aperture: clients and backends placed at equal intervals on two rings, P2C within the window, fractional overlap handled explicitly.
Design ruleJudge a subsetting scheme by its worst backend, not its mean. p99/average load ratio (Uber's metric) is the number to put on the dashboard before and after.
Case study

Kubernetes + gRPC: the balancer that balanced nothing

AssumptionA Kubernetes Service in front of N pods spreads the requests across them.
What happenedkube-proxy balances connections, and "HTTP/2 is designed to have a single long-lived TCP connection, across which all requests are multiplexed". The demo app's requests all rode one connection, so "only one of the pods is receiving any traffic".
Blast radiusA structural ceiling rather than an incident: one pod's capacity, regardless of replica count. Widely rediscovered by gRPC adopters since 2018.
FixRequest-level balancing: an L7 proxy or sidecar (the post demonstrates Linkerd), or a client-side gRPC balancer holding a connection per backend.
Design ruleAsk what unit your balancer balances. If the protocol multiplexes, a connection balancer is a pinning machine, and adding replicas changes nothing.

Figure 5 · One replica's life, as the balancer sees it

discovery announces host

first eligible pick

slow-start window elapses

errors accumulate

outlier threshold crossed

passes active health check

discovery withdraws host

Registered

Warming

Active

Suspect

Ejected

Removed

discovery announces host

first eligible pick

slow-start window elapses

errors accumulate

outlier threshold crossed

passes active health check

discovery withdraws host

Registered

Warming

Active

Suspect

Ejected

Removed

The guards are the transitions, and each transition is where a failure class lives: joining (cold stampede), Suspect and back (sinkhole and flapping), the view ageing (Slack). States from the Envoy slow start and panic threshold docs.
Diagram source
05

Numbers you can plan against

Every figure carries its context and date. Defaults are design decisions someone made for you; the dates tell you how recently anyone reconsidered them.

MetricValueAtContextAs ofSource
Connection reduction from subsetting>99%TwitterRandom aperture vs full mesh, aggregate connections (measured)2019Twitter
Hot-instance load under random subsets~400 rpsTwitterBanding, while other instances sat near idle (measured)2019Twitter
Edge traffic during LB redesign>1M rpsNetflixStated scale motivating the choice-of-2 redesign (claimed by operator)2018Netflix
Probe rate~3 per queryGoogle (YouTube)Prequal's asynchronous probes, reusable across picks (measured)2024NSDI '24
Hot/cold RIF thresholdp80Google (YouTube)Replicas above the 80th percentile of recent requests-in-flight are avoided (design value)2024NSDI '24
Tail-latency gain from herd-aware selectionup to 3xC3 / Cassandrap99.9 improvement on EC2 benchmarks (measured, academic)2015NSDI '15
P2C default sample size2Envoyleast_request choice_count default; weights equal (design value)2026Envoy docs
Panic threshold50%Envoy, FinagleBelow this availability, health status is disregarded (default in both)2026Envoy docs
Trust blackout for new backends10 sgRPC (A58)Load reports ignored until reported continuously this long (default)2023gRFC A58
Weight expiry3 mingRPC (A58)Stale self-reports stop being used after this (default)2023gRFC A58
Slow-start floor weight10%EnvoyMinWeightPercent so a warming host still gets picked (default)2026Envoy docs
Queue-time alarm line40 msHerokuAverage New Relic queue time "indicative of a problem" under random routing2013Heroku FAQ
Stale pool view age at failure>8 hSlackHAProxy instances stuck with stale backend state before the outage (measured)2020Slack
Rebalance blast radius1.8% / 15 minGitHubActions runs delayed / average delay, one rebalancing event (measured)2026GitHub
Peak failure rate, retries vs balancer75%GitHubSearch queries failing while retry queues exhausted the LBs (measured)2025GitHub
Imbalance metricp99 / avg CPUUberThe ratio Uber defines as its load-imbalance KPI across a service's tasks2022Uber
Read these carefully

Netflix's 1M rps is the operator's own scale statement, not an audited measurement. Prequal's production gains ("much higher utilization", lower tails) are reported by its authors without company-external verification; the mechanism numbers (3 probes, p80) are from the paper and slides. Uber published its imbalance metric definition but the before and after ratios sit in figures this session could not extract; treat the metric as usable and the magnitude of Uber's win as unverified. C3's 3x is a benchmark result, not a production figure. Defaults (panic 50%, blackout 10 s) are current as of the dates shown and change with releases.

06

The evidence wall

Every source behind this page, graded. The strongest material is at the tiers ordinary search never surfaces: the postmortems, the design records, and one closed, unmerged proposal.

Postmortem Slack2020-05

A Terrible, Horrible, No-Good, Very Bad Day at Slack

The definitive pool-view-drift incident: HAProxy state over eight hours stale, autoscaling removing exactly the instances the stale view still trusted, and the one alert that would have caught it silently broken.

Carry forwardAlert on the age of the balancer's view, not only on backend health.
slack.engineering
Postmortem GitHub2026-02

Availability report: February 2026

A connection rebalancing event inside the internal LB layer skewed traffic across sites and triggered throttling: 1.8% of Actions runs delayed by 15 minutes on average, from housekeeping, not from failure.

Carry forwardLB reloads synchronize client reconnects; spread them deliberately.
github.blog
Postmortem GitHub2025-08

Availability report: August 2025

Retries first masked flaky LB-to-search connectivity, then exhausted the balancers themselves; up to 75% of search queries failed until the indexing pipeline was slowed.

Carry forwardRetry budgets must count the balancer's state, not just backend capacity.
github.blog
Postmortem GitHub2024-05

Availability report: May 2024

A provider-side OS upgrade produced "unintended and uneven traffic distribution within the cluster"; the remediation list included closing monitoring gaps for load thresholds.

Carry forwardWatch the distribution itself; even a managed layer can skew silently.
github.blog
Postmortem Heroku2013-02

Routing Performance Update

The platform's official admission that routing behaviour had drifted from its documentation and caused years of unexplained Rails latency; paired with a FAQ that defends random routing and names its precondition.

Carry forwardRandom is fine for concurrent backends and poison for serial ones.
heroku.com
Source gRPC2023–24

proposal PR #383: A68 deterministic subsetting, closed unmerged

The recorded argument: a full xDS policy for SRE-book-style deterministic subsetting, drafted, implemented in Go and Java, and closed without merging, with the doc still marked Draft.

Carry forwardDeterministic subsetting's fairness was not worth its churn and complexity to gRPC.
github.com/grpc/proposal
Source gRPC2024–26

proposal PR #423: A68 random subsetting with rendezvous hashing

What won instead: per-process salted rendezvous hashing over the address list, shipped as the experimental randomsubsetting balancer in grpc-go.

Carry forwardSalted randomness decorrelates clients; determinism correlates them on purpose. Pick which correlation you want.
github.com/grpc/proposal
Source Envoychecked 2026-10

Issue #17013: deterministic aperture in Envoy

The feature request to port Twitter's d-aperture, with the benefit argued in-thread (fewer connections, even spread, no coordination) and no implementation to date.

Carry forwardThe gap between published algorithm and available implementation is measured in years; plan around what your proxy ships.
github.com/envoyproxy
ADR gRPC2023-04

gRFC A58: weighted_round_robin LB policy

The most explicit trust contract in the corpus: weight = qps / (utilization + eps/qps × penalty), a 10 s blackout before a backend's reports count, 3 min expiry after they stop, and an error term so failure cannot masquerade as capacity.

Carry forwardA self-reported signal is usable exactly to the extent you bound when it may be believed.
github.com/grpc/proposal
ADR gRPC2023–24

A68 design doc (deterministic subsetting draft)

The draft design that carried Google's SRE-book subsetting algorithm into xDS configuration, including the leftover-task refinement; valuable precisely because it was not accepted.

Carry forwardRead rejected designs for the constraint list the accepted one had to beat.
github.com/grpc/proposal
Case study Google2016

SRE book, ch. 20: Load Balancing in the Datacenter

The chapter that documents the error sinkhole ("frequently significantly faster to just return an 'I'm unhealthy!' error than to actually process a request"), deterministic subsetting, and weighted round robin on backend-reported load.

Carry forwardAny busy-ness score needs an error term, or failure reads as capacity.
sre.google
Blog Netflix2018-09

Rethinking Netflix's Edge Load Balancing

Round robin plus blacklisting was not enough at 1M+ rps; the replacement combines choice-of-2 with the balancer's view first and the server's self-report second, with adaptive guardrails instead of static thresholds.

Carry forwardClient sees latency best; server sees its own utilization best. Use both, in that order.
netflixtechblog.com
Blog Twitter2019

Deterministic Aperture

The three-act subsetting story with production numbers: full mesh unaffordable, random subsets cut connections 99% but banded the load, ring-based determinism restored fairness.

Carry forwardSubsetting decisions show up in the worst replica's load, so measure that.
blog.twitter.com
Blog Uber2022-05

Better Load Balancing: Real-Time Dynamic Subsetting

Subset sizes recomputed from aggregated real-time load reports across a mesh of thousands of services in millions of containers; defines the p99/average CPU imbalance metric this guide recommends.

Carry forwardA fixed subset size is wrong for someone; size it from observed load.
eng.uber.com
Blog Uber2024-03

Load Balancing: Handling Heterogeneous Hardware

The year-long follow-up: weighting hosts by hardware generation, framed and funded as efficiency work rather than reliability work.

Carry forward"Identical replicas" is a fiction with a hardware refresh cycle; weights must encode it.
uber.com
Blog AWS (M. Brooker)2012 / 2024

Two random choices; Best-of-K

Twelve years apart, the same conclusion from simulation and large-scale deployment: cached global state herds, best-of-k with k of 2 or 3 is nearly as good as perfect information and far more robust to staleness.

Carry forwardStaleness tolerance, not optimality, is the property to buy.
brooker.co.za
Blog Buoyant / Kubernetes2018-11

gRPC Load Balancing on Kubernetes without Tears

The connection-versus-request unit mismatch demonstrated on a live app: one pod took all traffic because HTTP/2 multiplexes onto one long-lived connection.

Carry forwardName the unit your balancer balances before tuning its algorithm.
kubernetes.io
Blog Linkerd2016-03

Beyond Round Robin: Load Balancing for Latency

The early public comparison showing latency-aware policies (least-loaded, peak EWMA) beating the round robin that nginx and HAProxy made the industry default.

Carry forwardThe default algorithm in your proxy is an accident of history, not a recommendation.
linkerd.io
Blog Heroku2013-02

Routing and Web Performance on Heroku: a FAQ

The defence of random routing on availability grounds, the 40 ms queue-time alarm line, and the precise description of which applications the policy hurt.

Carry forwardPublish your balancer's contract; Heroku's outage was partly a documentation failure.
heroku.com
Paper Mitzenmacher2001

The Power of Two Choices in Randomized Load Balancing

The supermarket-model result the whole field leans on: d=2 gives an exponential improvement over random, d=3 only a constant factor more. The maths that makes sampling two sufficient.

Carry forwardBuy the second choice; the third is not worth its coordination cost.
eecs.harvard.edu
Paper Google2024-04

Load is not what you should balance: Introducing Prequal (NSDI '24)

YouTube's balancer: asynchronous reusable probes (~3 per query), hot/cold classification at the 80th percentile of requests-in-flight, lowest latency among the cold. CPU is deliberately not the signal.

Carry forwardBalance what predicts waiting (queue, latency), not what bills (CPU).
usenix.org
Paper TU Berlin / UCLouvain2015-05

C3: Adaptive Replica Selection (NSDI '15)

Replica selection inside a data store, with ranking built to avoid herding onto fast servers with growing queues; up to 3x p99.9 improvement over Cassandra's stock selection.

Carry forwardPenalise queue depth superlinearly or your fastest replica becomes your hottest.
usenix.org
Paper Google2023-05

Reinventing Backend Subsetting at Google (CACM)

A decade after the SRE book published deterministic subsetting, Google wrote up its replacement, designed around reducing connection churn during rollouts.

Carry forwardSubsetting quality is churn under change, not just spread at rest.
cacm.acm.org
Talk Fastly2016-11

Load Balancing is Impossible (Tyler McMullen, QCon SF)

The talk that names the physics: Poisson arrivals and heavy-tailed service times mean perfect balance is unattainable, and the practical frontier is randomized least-conns, Join-Idle-Queue and careful load interpretation. Cited to the talk page and slide deck; this build environment could not play the video to timestamp claims.

Carry forwardAim for bounded worst-case imbalance, not perfect balance.
infoq.com
Talk Twitter / USENIX2019-03

Aperture: A Non-Cooperative, Client-Side Load Balancing Algorithm (SREcon19)

The conference version of the aperture story, explicit that thousands of independent client-side balancers must reach fairness without coordinating. Cited to the talk page and the SREcon APAC slide deck; video not viewable from this build environment.

Carry forwardWith client-side balancing, fairness is an emergent property you must design for, not configure.
usenix.org
Vendor Envoy2026

Supported load balancers (architecture docs)

The least-request policy's two modes, the P2C default with the Mitzenmacher citation in the docs themselves, and the admission that the weighted fallback "will never truly drain" a host.

Carry forwardCheck which mode your weights push you into; they are different algorithms.
envoyproxy.io
Vendor Envoy2026

Slow start mode

The cold-host guard and its own failure modes, documented: ineffective when the whole fleet is new, risky at low traffic (starvation, non-gradual weight jumps).

Carry forwardSlow start protects against a few cold hosts, not a cold fleet; deploys need their own ramp.
envoyproxy.io
Vendor Envoy2026

Panic threshold

Below 50% availability Envoy stops trusting health entirely, balancing to all hosts or failing outright, with the trade-off between the two modes spelled out.

Carry forwardDecide before the incident whether a dying pool should spread load or shed it.
envoyproxy.io
Vendor Twitter (Finagle)2026

Finagle client docs: load balancing (Clients.rst)

Unusually honest library documentation: the heap balancer's contention, P2C's degradation to random under low load, Peak EWMA's long-polling blind spot, panic mode's thresholds.

Carry forwardEvery load metric has a traffic pattern that blinds it; the docs of a 15-year-old balancer list them.
github.com/twitter/finagle
Vendor AWS2019-11

ALB adds Least Outstanding Requests

The managed-platform data point: the biggest cloud load balancer offered only round robin until late 2019, and added LOR citing varied request costs and target churn.

Carry forwardManaged defaults lag the state of the art by years; check what your LB actually runs.
aws.amazon.com
07

Build a miniature, then productionise it

Six rungs. The first three fit in an evening each with two toy HTTP servers and any proxy; the last three are the production crossing.

Feel the lumpiness of random

Two identical backends, a proxy assigning uniformly at random, constant request rate. Plot per-backend in-flight counts over ten minutes, then repeat with round robin.

Done when: your plot shows random producing transient 3:1 in-flight skews that round robin does not.  Teaches: why "uniform" and "even" are different words, the gap behind Heroku's 2013 trouble.

Build the error sinkhole, on purpose

Make one backend return 500 in 2 ms while the other serves in 120 ms. Switch the proxy to least-connections and watch the traffic share of the broken backend.

Done when: the failing backend receives the clear majority of requests within a minute.  Teaches: the SRE book's sinkhole, viscerally; why every busy-ness score needs an error term.

Fix it with P2C plus an error penalty

Implement pick-two-compare on in-flight count, then multiply each candidate's score by a penalty from its recent error rate (gRFC A58's eps/qps shape). Re-run rung 2.

Done when: the broken backend's share falls below 5% without any health check marking it down.  Teaches: the difference between health checking and selection, and why both exist.

Cross into production shape: guards in a real proxy

Run Envoy with least_request, slow start (30 s window) and outlier ejection in front of three backends. Kill and cold-start one backend repeatedly while under load.

Done when: a restarting backend's traffic ramps visibly in the Envoy stats and p99 stays flat through restarts.  Teaches: guards as config you can read, and what the stats for them look like on a dashboard.

Make a backend lie to you

Add a self-reported utilization value to each backend's responses and switch selection to weight on it. Then freeze one backend's reports and make another under-report by half.

Done when: you have reproduced both failure modes and fixed them with a blackout period and an expiry, matching A58's 10 s / 3 min defaults.  Teaches: why a signal contract has three clauses, not one.

Subset and measure the worst replica

Fifty client processes, twenty backends, subset size five. Compare random subsets against deterministic ring windows. Track connection count and p99/average load per backend (Uber's metric), then roll all backends and measure connection churn.

Done when: you can state the three-way trade you measured: connections saved, worst-backend load, churn under rollout.  Teaches: why Twitter, Google and gRPC each landed on a different subsetting answer.

08

Keep hunting

The queries that actually found this material, in the forms that worked. The vocabulary terms (aperture, banding, requests-in-flight, p99/average) are the keys; each good source taught the next query.

Incidents and postmortems

  • "least connections" load balancer failing server outage postmortem
  • "connection rebalancing" availability report site:github.blog
  • slack outage 2020 haproxy "server state" postmortem
  • heroku "routing performance update" random routing

Production accounts and vocabulary

  • "deterministic aperture" load balancing
  • uber engineering "dynamic subsetting"
  • "p99" "average cpu" imbalance load balancing
  • prequal "requests in flight" nsdi youtube

Design records and rejected approaches

  • repo:grpc/proposal is:pr is:closed is:unmerged subsetting
  • grpc proposal "weighted_round_robin" blackout_period
  • envoy issue "deterministic aperture"
  • "reinventing backend subsetting" google

Mechanism and theory

  • mitzenmacher "power of two choices" supermarket model
  • brooker "best of k" stale load balancing
  • "herd behavior" replica selection tail latency
  • envoy "slow start" starvation "low traffic"
09

References

  1. Google, Site Reliability Engineering, ch. 20: Load Balancing in the Datacenter O'Reilly / sre.google, 2016. Checked 2026-10-09.
  2. Mike Smith, Rethinking Netflix's Edge Load Balancing Netflix Technology Blog, 2018-09-28. Checked 2026-10-09.
  3. Ruben Oanta & Bryce Anderson, Deterministic Aperture Twitter Engineering, 2019. Checked 2026-10-09.
  4. Aperture: A Non-Cooperative, Client-Side Load Balancing Algorithm SREcon19 Americas, USENIX, 2019-03. Checked 2026-10-09. Slides: SREcon19 APAC deck.
  5. Liao, Kundu & Krolikowski, Better Load Balancing: Real-Time Dynamic Subsetting Uber Engineering, 2022-05-17. Checked 2026-10-09.
  6. Jiang, Liao & Krolikowski, Load Balancing: Handling Heterogeneous Hardware Uber Engineering, 2024-03-07. Checked 2026-10-09.
  7. Laura Nolan, A Terrible, Horrible, No-Good, Very Bad Day at Slack Slack Engineering, 2020. Checked 2026-10-09.
  8. Slack status: incident of 2020-05-12 Slack, 2020-05-12. Checked 2026-10-09.
  9. GitHub Availability Report: February 2026 GitHub, 2026-03. Checked 2026-10-09.
  10. GitHub Availability Report: August 2025 GitHub, 2025-09. Checked 2026-10-09.
  11. GitHub Availability Report: May 2024 GitHub, 2024-06-12. Checked 2026-10-09.
  12. Heroku, Routing Performance Update Heroku, 2013-02. Checked 2026-10-09.
  13. Heroku, Routing and Web Performance on Heroku: a FAQ Heroku, 2013-02. Checked 2026-10-09.
  14. The Register, Heroku marketing misstep draws customer anger The Register, 2013-02-15. Checked 2026-10-09.
  15. Marc Brooker, The power of two random choices brooker.co.za, 2012-01-17. Checked 2026-10-09.
  16. Marc Brooker, Finding Needles in a Haystack with Best-of-K brooker.co.za, 2024-03-25. Checked 2026-10-09.
  17. Mitzenmacher, The Power of Two Choices in Randomized Load Balancing IEEE TPDS, 2001. Checked 2026-10-09.
  18. Wydrowski, Kleinberg, Rumble & Archer, Load is not what you should balance: Introducing Prequal USENIX NSDI, 2024-04. Checked 2026-10-09.
  19. Suresh, Canini, Schmid & Feldmann, C3: Cutting Tail Latency in Cloud Data Stores via Adaptive Replica Selection USENIX NSDI, 2015-05. Checked 2026-10-09.
  20. Reinventing Backend Subsetting at Google Communications of the ACM, 2023-05. Checked 2026-10-09.
  21. gRFC A58: weighted_round_robin LB policy gRPC proposal repo, 2023-04-17. Checked 2026-10-09 (fetched in full).
  22. gRPC proposal PR #383: A68 Deterministic Subsetting LB policy (closed unmerged) GitHub, 2023–2024. Checked 2026-10-09.
  23. gRPC proposal PR #423: A68 Random subsetting with rendezvous hashing GitHub, 2024–2026. Checked 2026-10-09.
  24. Envoy issue #17013: deterministic aperture GitHub, checked 2026-10-09.
  25. Envoy architecture: supported load balancers Envoy docs; source fetched from envoyproxy/envoy main, 2026-10-09.
  26. Envoy architecture: slow start mode Envoy docs; source fetched from envoyproxy/envoy main, 2026-10-09.
  27. Envoy architecture: panic threshold Envoy docs; source fetched from envoyproxy/envoy main, 2026-10-09.
  28. Envoy contrib: Peak EWMA load balancer Envoy docs. Checked 2026-10-09.
  29. Finagle client documentation: load balancing twitter/finagle repo, develop branch; fetched in full 2026-10-09.
  30. William Morgan, gRPC Load Balancing on Kubernetes without Tears Kubernetes blog, 2018-11-07. Checked 2026-10-09.
  31. Linkerd, Beyond Round Robin: Load Balancing for Latency linkerd.io, 2016-03-16. Checked 2026-10-09.
  32. AWS, ALB now supports Least Outstanding Requests AWS What's New, 2019-11-25. Checked 2026-10-09.
  33. Tyler McMullen, Load Balancing is Impossible QCon SF 2016 via InfoQ. Checked 2026-10-09. Slides: QCon deck.