Evidence ledger 30 sources Checked 09 Oct 2026

Evidence ledger

One row per claim in Choosing the backend: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

One row per claim. Tiers follow the skill's grading: postmortem, source, adr, casestudy, blog, paper, talk, vendor. All links checked 2026-10-09. Network note: this session ran in a container whose egress policy allowed only GitHub and package hosts, so GitHub-hosted sources (gRFCs, repo docs, issues/PRs) were fetched raw and in full; all other pages were retrieved and quoted through the session's search tool, which fetches pages server-side. No claim below is cited from memory; every quote was copied from content retrieved this session.

# Org Title Tier Published Checked URL Claim I take from it Supporting quote or figure
1 Google SRE book, ch. 20: Load Balancing in the Datacenter casestudy 2016 2026-10-09 https://sre.google/sre-book/load-balancing-datacenter/ Least-loaded routing sends more traffic to a backend that is failing fast, because errors are cheaper than work "If a task is seriously unhealthy, it might start serving 100% errors. Depending on the nature of those errors, they may have very low latency; it's frequently significantly faster to just return an 'I'm unhealthy!' error than to actually process a request." Clients then send "a very large amount of traffic to the unhealthy task"
2 Google SRE book, ch. 20 casestudy 2016 2026-10-09 https://sre.google/sre-book/load-balancing-datacenter/ Google runs deterministic subsetting so each client connects to a fraction of backends; the chapter's algorithm divides clients into rounds, one backend per client per round Chapter describes dividing "client tasks into rounds", each backend "assigned to one client within a round" (algorithm reimplemented verbatim in [23])
3 Netflix Rethinking Netflix's Edge Load Balancing (Mike Smith) blog 2018-09-28 2026-10-09 https://netflixtechblog.com/rethinking-netflixs-edge-load-balancing-695308b5548c The conditions that overload single servers at the edge: cold servers after deploy/autoscale, GC pauses, permanently slower hardware; motivation was cutting load-related errors at >1M req/s Post names "cold servers right after startup during deployments and autoscaling, servers briefly stalling during configuration updates or large garbage-collection pauses, and hardware differences"; "even a low rate of errors at our scale of over a million requests per second can degrade the experience for our members"
4 Netflix Rethinking Netflix's Edge Load Balancing blog 2018-09-28 2026-10-09 https://netflixtechblog.com/rethinking-netflixs-edge-load-balancing-695308b5548c The replacement combines a choice-of-2 pick with two signals: the balancer's view (in-flight count) primarily and the server's self-reported utilization secondarily; guardrails are adaptive rather than static thresholds "client-side views are the best source for a server's latency, while the server itself is the best source for its utilization"; balancer picks "primarily on the load balancers' view of a server's utilization, and secondarily on the servers' view"; "instead of configuring static thresholds, they use adaptive mechanisms that change based on current traffic, performance, and environment"
5 Twitter Deterministic Aperture: A distributed, load balancing algorithm (Oanta/Anderson) blog 2019 2026-10-09 https://blog.twitter.com/engineering/en_us/topics/infrastructure/2019/daperture-load-balancer Random subsetting cut aggregate connection count by over 99% but produced banded, uneven request distribution with hotspots near 400 rps "While we succeeded in dramatically reducing our aggregate connection count by over 99%, we've made a mess of our request distribution. Interestingly, we can see the formation of some distinct bands of load" (hot instances approaching 400 requests per second)
6 Twitter Deterministic Aperture blog 2019 2026-10-09 https://blog.twitter.com/engineering/en_us/topics/infrastructure/2019/daperture-load-balancer Deterministic aperture places clients and backends on two rings at equal intervals and runs P2C inside the window, keeping the connection savings while restoring fairness Clients and backends "are positioned on these rings at equidistant intervals" (peer ring / destination ring); subsetting "was the source of the gains, while doing it randomly was the source of the problems"
7 Uber Better Load Balancing: Real-Time Dynamic Subsetting blog 2022-05-17 2026-10-09 https://eng.uber.com/better-load-balancing-real-time-dynamic-subsetting/ Uber's mesh (on-host L7 proxy, RPC between thousands of services in millions of containers) measures balance as p99/average CPU and sizes subsets from real-time aggregated load reports "CPU load imbalance is defined as the ratio of p99 to average CPU utilization of tasks for a given service"; mesh has "been powering RPC traffic between thousands of microservices since 2016" "(in millions of containers)"
8 Uber Load Balancing: Handling Heterogeneous Hardware blog 2024-03-07 2026-10-09 https://www.uber.com/blog/load-balancing-handling-heterogeneous-hardware/ The follow-up work to weight hosts by hardware generation ran over a year across multiple teams and is framed as efficiency work Work "lasted over a year, involved engineers across multiple teams, and delivered significant efficiency savings"
9 Marc Brooker (AWS) The power of two random choices blog 2012-01-17 2026-10-09 https://brooker.co.za/blog/2012/01/17/two-random/ Balancing from a cached global load snapshot is tempting and wrong; two random choices against live local state beats it "The overhead of constantly sharing the exact load information between different sources can be high, so it's tempting to have each source work off a cached copy"; "turns out that's not a great idea"
10 Marc Brooker (AWS) Finding Needles in a Haystack with Best-of-K blog 2024-03-25 2026-10-09 https://brooker.co.za/blog/2024/03/25/needles/ Best-of-k (k≈2..3) is much more robust to stale data than best-of-n and O(1); author has deployed it repeatedly in large production systems; exception is very small bins Snapshot approaches work "in slow-moving systems, but the stale data quickly causes bad decisions"; best-of-k "is much more robust to stale data than best-of-n"; "deployed them many times in large-scale production systems"
11 Harvard / Mitzenmacher The Power of Two Choices in Randomized Load Balancing (IEEE TPDS 2001) paper 2001 2026-10-09 https://www.eecs.harvard.edu/~michaelm/abstracts/tpds2001.html d=2 choices gives an exponential improvement over d=1 in expected time in system; d=3 is only a constant factor better than d=2 "having d=2 choices leads to exponential improvements in the expected time a customer spends in the system over d=1, whereas having d=3 choices is only a constant factor better than d=2" (supermarket model)
12 Google Load is not what you should balance: Introducing Prequal (NSDI '24) paper 2024-04 2026-10-09 https://www.usenix.org/conference/nsdi24/presentation/wydrowski Google's YouTube serving balancer deliberately does not balance CPU load; it selects on estimated latency and requests-in-flight via asynchronous probes "Cutting against received wisdom, Prequal does not balance CPU load, but instead selects servers according to estimated latency and active requests-in-flight (RIF)"; probes run at "about 3 probes per query"; hot/cold rule: RIF above the 80th percentile of recent RIF values = hot, avoided; among cold, pick lowest latency
13 Google Prequal (NSDI '24) paper 2024-04 2026-10-09 https://www.usenix.org/conference/nsdi24/presentation/wydrowski Production result claim: deployed on YouTube for over a year, with large reductions in tail latency, errors and resource use at higher utilization Abstract: Prequal "has dramatically decreased tail latency, error rates, and resource use, enabling YouTube and other production systems at Google to run at much higher utilization"
14 TU Berlin / UCLouvain C3: Cutting Tail Latency in Cloud Data Stores via Adaptive Replica Selection (NSDI '15) paper 2015-05 2026-10-09 https://www.usenix.org/conference/nsdi15/technical-sessions/presentation/suresh Replica selection under fluctuating server performance is its own problem; adaptive selection with herd-avoidance cut Cassandra p99.9 latency up to 3x Abstract: "C3 significantly improves the latencies along the mean, median, and tail (up to 3 times improvement at the 99.9th percentile) and provides higher system throughput"; ranking deliberately ranks fast servers with long queues low "to avoid herd behavior" (course scribe of the paper)
15 Google Reinventing Backend Subsetting at Google (CACM, May 2023) paper 2023-05 2026-10-09 https://cacm.acm.org/magazines/2023/5/272290-reinventing-backend-subsetting-at-google/fulltext Google replaced deterministic subsetting because of connection churn during rollouts; subsetting algorithm choice is still active research a decade after the SRE book Article abstract: designing an algorithm that "reduces connection churn and could replace deterministic subsetting"
16 gRPC gRFC A58: weighted_round_robin LB policy adr 2023-04-17 2026-10-09 https://github.com/grpc/proposal/blob/master/A58-client-side-weighted-round-robin-lb-policy.md The design's weight formula and its three trust guards: blackout before trusting a new backend's reports, expiry of stale weights, error penalty on the utilization signal weight = qps / (utilization + eps/qps × error_utilization_penalty); blackout_period "Default is 10 seconds"; weight_expiration_period "ensures that we do not continue to use very stale weights... Defaults to 3 minutes"; fetched raw in full
17 gRPC proposal PR #383: A68 Deterministic Subsetting LB policy (closed unmerged) source closed 2024 (opened 2023) 2026-10-09 https://github.com/grpc/proposal/pull/383 gRPC considered deterministic subsetting as a first-class policy and closed the proposal without merging; the design doc never left Draft status PR titled "A68: Deterministic Subsetting LB policy", state Closed; doc status line "Draft"
18 gRPC proposal PR #423: A68 Random subsetting with rendezvous hashing source 2024–2026 2026-10-09 https://github.com/grpc/proposal/pull/423 What replaced it: random subsetting with rendezvous hashing; shipped as experimental randomsubsetting balancer in grpc-go v1.80 PR text: "when the lb policy is initialized it also creates a random 32-byte long salt string"; godocs lists google.golang.org/grpc/balancer/randomsubsetting as experimental
19 Envoy Supported load balancers: weighted least request vendor current docs (fetched 2026-10-09 from repo main) 2026-10-09 https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/load_balancers Envoy's least-request is P2C when weights are equal, citing Mitzenmacher, chosen for "resistance to herding behavior"; with unequal weights it degrades to a dynamic WRR that never truly drains a host "An O(1) algorithm which selects N random available hosts... (2 by default) and picks the host which has the fewest active requests"; "P2C selection is particularly useful... due to its resistance to herding behavior"; "unlike P2C, a host will never truly drain"
20 Envoy Slow start mode vendor current docs (fetched 2026-10-09 from repo main) 2026-10-09 https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/slow_start Without slow start a new endpoint gets a full proportional share immediately; slow start is useless when the whole fleet is new, and harmful at low traffic or high endpoint counts "With no slow start enabled Envoy would send a proportional amount of traffic to new upstream endpoints... could result in request timeouts, loss of data"; "When all the endpoints are relatively new e.g. new deployment in Kubernetes, slow start is not very effective"; "not recommended... in low traffic or high number of endpoints scenarios" (starvation, non-gradual jumps)
21 Envoy Panic threshold vendor current docs (fetched 2026-10-09 from repo main) 2026-10-09 https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/panic_threshold Below 50% available hosts (default) Envoy stops honouring health and balances across all hosts, or none, to stop failures cascading "if the percentage of available hosts in the cluster becomes too low, Envoy will disregard health status and balance either amongst all hosts or no hosts... The default panic threshold is 50%"
22 Envoy issue #17013: deterministic aperture request source open (checked 2026-10-09) 2026-10-09 https://github.com/envoyproxy/envoy/issues/17013 Twitter's d-aperture has been requested in Envoy and discussed, not implemented; the thread records the claimed benefit Comment: "the d-aperture balancer reduces the total number of connections used by N clients communicating with M upstream servers, while spreading requests evenly without requiring coordination across clients"
23 gRPC proposal PR #383 review thread / SRE-book subsetting reimplementation source 2023–2024 2026-10-09 https://github.com/grpc/proposal/pull/383/files The A68 draft carried the SRE book's deterministic subsetting algorithm into gRPC xDS config, with the known leftover-task refinement Doc describes a deterministic_subsetting policy and the Chapter 20 algorithm "with one improvement: balancing the leftover tasks in each group... in a round-robin fashion"
24 Twitter Finagle client documentation (Clients.rst, develop branch) vendor fetched raw in full 2026-10-09 2026-10-09 https://github.com/twitter/finagle/blob/develop/doc/src/sphinx/Clients.rst Finagle's defaults and caveats: P2C+least-loaded is the default; least-loaded degrades toward random without concurrent load; Peak EWMA penalises slow endpoints but breaks with long-polling; panic mode gives up finding a healthy node past a threshold "without sufficient concurrent load, the previous distributors can degrade to random selection"; Peak EWMA "is designed to react to slow endpoints more quickly than least loaded... this assumption breaks down in the presence of long polling clients"; "Panic mode is when the load balancer gives up trying to find a healthy node" (default FiftyPercentUnhealthy)
25 Slack A Terrible, Horrible, No-Good, Very Bad Day at Slack (Laura Nolan) postmortem 2020 2026-10-09 https://slack.engineering/a-terrible-horrible-no-good-very-bad-day-at-slack/ May 12 2020 outage: HAProxy instances held stale backend state (most >8h old); autoscaling removed oldest instances first so the state pointed at dead hosts; the alert for exactly this was broken; fix was a rolling restart "Most of them were more than eight hours old and therefore were stuck with full and stale backend state"; "there were no longer enough older webapp instances remaining in the HAProxy server state to serve demand"; "We had alerting in place for this precise situation, but unfortunately, it wasn't working as intended"
26 Slack Status page incident, 2020-05-12 postmortem 2020-05-12 2026-10-09 https://slack-status.com/2020-05/147dad376c8946ff The user-facing account: servers failed to register with the load balancer during scale-up; pool health declined over time "Some of these servers did not successfully register with our load balancing infrastructure during this process of scaling up, and this ultimately led to a decline in the health of the server pool over time" (outage window 4:45–5:33 p.m. PDT)
27 GitHub Availability report: February 2026 (incident of Feb 23) postmortem 2026-03 2026-10-09 https://github.blog/news-insights/company-news/github-availability-report-february-2026/ One connection rebalancing event inside the internal LB layer skewed traffic across sites and throttled requests; 1.8% of Actions runs delayed, average 15 minutes "a connection rebalancing event in our internal load balancing layer, which temporarily created uneven traffic distribution across sites and led to request throttling"; "1.8% of Actions workflow runs experienced delayed starts with an average delay of 15 minutes" (15:00–17:00 UTC)
28 GitHub Availability report: August 2025 (incident of Aug 12) postmortem 2025-09 2026-10-09 https://github.blog/news-insights/company-news/github-availability-report-august-2025/ Retry logic first hid intermittent LB-to-search-host connectivity loss, then the accumulated retries exhausted the load balancers themselves; up to 75% of search queries failed Degradation 13:30–17:14 UTC; "up to 75% of search queries failed"; retries "hid the problem at first, but then the accumulated retries exhausted the load balancers and they failed"
29 GitHub Availability report: May 2024 postmortem 2024-06-12 2026-10-09 https://github.blog/2024-06-12-github-availability-report-may-2024 A provider-side OS upgrade caused "unintended and uneven traffic distribution" in a cluster; GitHub's fix included monitoring gaps for load thresholds May 21 incident: "a scheduled operating system upgrade that led to unintended and uneven traffic distribution within the cluster"; mitigation added network routes and fixed "gaps in monitoring and alerting for load thresholds"
30 Heroku Routing Performance Update postmortem 2013-02 2026-10-09 https://www.heroku.com/blog/routing_performance_update/ Heroku's official admission that routing on Bamboo/Cedar caused unexplained latency for Rails apps and that documentation misdescribed the routing behaviour Post admits routing caused "unexplained high latencies, mismatched queuing metrics, and differences between documented and observed behavior", and that Heroku "failed to properly document" Bamboo routing
31 Heroku Routing and Web Performance on Heroku: a FAQ blog 2013-02 2026-10-09 https://www.heroku.com/blog/routing_and_web_performance_on_heroku_a_faq/ Heroku's stated reason for random routing over a global request queue: availability and stateless horizontal scaling; the affected population and the 40ms queue-time signal "The Heroku router favors availability, stateless horizontal scaling, and low latency through individual routing nodes"; found "no model or implementation that beat the simplicity and robustness of random routing to back-ends that support multiple concurrent connections"; "average queue times above 40ms are usually indicative of a problem"; most affected: "Rails apps running on Thin, with six or more dynos, and serving 1k requests per minute or more"
32 Kubernetes / Buoyant gRPC Load Balancing on Kubernetes without Tears (William Morgan) blog 2018-11-07 2026-10-09 https://kubernetes.io/blog/2018/11/07/grpc-load-balancing-on-kubernetes-without-tears/ HTTP/2 multiplexes everything over one long-lived connection, so Kubernetes' connection-level balancing pins all gRPC requests to one pod "HTTP/2 is designed to have a single long-lived TCP connection, across which all requests are multiplexed"; result in the demo app: "only one of the pods is receiving any traffic"; fix: "open an HTTP/2 connection to each destination, and balance requests across these connections"
33 Linkerd Beyond Round Robin: Load Balancing for Latency blog 2016-03-16 2026-10-09 https://linkerd.io/2016/03/16/beyond-round-robin-load-balancing-for-latency/ Round robin is the common default in nginx/HAProxy; latency-aware policies (least-loaded, peak EWMA) measurably beat it in Linkerd's comparison Round robin is "commonly seen in practice, and is available in most software load balancers, including Nginx and HAProxy"
34 AWS ALB now supports Least Outstanding Requests vendor 2019-11-25 2026-10-09 https://aws.amazon.com/about-aws/whats-new/2019/11/application-load-balancer-now-supports-least-outstanding-requests-algorithm-for-load-balancing-requests/ ALB ran round robin only until Nov 2019; AWS's stated reason for adding LOR is over/under-utilization with varied request costs or churn Round robin "led to over-utilization or under-utilization of targets when requests had varied processing times or targets were frequently added or removed"
35 Fastly / QCon Load Balancing is Impossible (Tyler McMullen, QCon SF 2016) talk 2016-11 2026-10-09 https://www.infoq.com/presentations/load-balancing/ The talk's thesis and the named techniques: perfect balancing is unattainable (Poisson arrivals, heavy-tailed service times); randomized least-conns, Join-Idle-Queue, load interpretation are the practical tools. Claims cited to the talk page and the published slide deck, not to video timestamps (video not viewable from this build environment) InfoQ summary: "discusses load balancing techniques and algorithms such as Randomized Least-conns, Join-Idle-Queue, and Load Interpretation"; abstract: "load balancing perfectly may be impossible in the real world, but we can do better than 'random', 'round-robin', and naive 'least-conns'"; slides present random as "the inglorious default" (deck: https://qconsf.com/sf2016/system/files/presentation-slides/load_balancing_is_impossible_-_qcon_sf_2016_.pdf)
36 Twitter / USENIX Aperture: A Non-Cooperative, Client-Side Load Balancing Algorithm (SREcon19 Americas, Oanta & Anderson) talk 2019-03 2026-10-09 https://www.usenix.org/conference/srecon19americas/presentation/oanta The conference account of aperture: non-cooperative client-side balancing, three-generation framing (P2C fair but costly at scale; random aperture scalable but unfair; deterministic aperture both). Cited to the talk page and the SREcon APAC slide deck, not to video timestamps (video not viewable from this build environment) Talk page: "Finagle uses non-cooperative, client-side load balancing"; APAC slides frame the progression: fair P2C → scalable-but-unfair random aperture → scalable and fair deterministic aperture (deck: https://www.usenix.net/sites/default/files/conference/protected-files/srecon19apac_slides_anderson.pdf)
37 Envoy Peak EWMA load balancer (contrib) vendor current docs 2026-10-09 https://www.envoyproxy.io/docs/envoy/latest/api-v3/config/contrib/load_balancing_policies/peak_ewma/peak_ewma Latency-aware selection reached Envoy only as a build-time contrib extension, and its docs warn it does not handle unhealthy hosts or error responses by itself Docs: contrib extension "must be explicitly enabled at build time"; "does not handle unhealthy hosts or error responses directly"
38 The Register Heroku marketing misstep draws customer anger blog 2013-02-15 2026-10-09 https://www.theregister.co.uk/2013/02/15/heroku_marketing_misstep_draws_customer_anger/ Independent press account: Heroku moved from "intelligent routing" to random routing around mid-2010 while continuing to describe it as intelligent The Register "dates the switch to mid-2010" and notes Heroku "kept using the 'intelligent routing' term even though requests were routed randomly"

Notes on gaps and grading

  • No public postmortem in this corpus names the least-loaded error sinkhole as the root cause of a specific dated incident. The failure is documented as a mechanism by Google's SRE book [1], and defences against it are built into gRFC A58 [16] (error penalty) and Netflix's design [4] (guardrails, server-reported health), and Envoy ships outlier ejection beside its least-request policy [19]. The page treats the mechanism as corroborated and the incident record as an open gap, and says so.
  • Talk citations carry no video timestamps. This build environment could not play or seek video; both talks [35][36] are cited to their published pages and slide decks instead, and the page says which claims rest on them.
  • Sources 25/26 are two accounts of one incident (engineering blog and status page) and are counted as one organisation's postmortem told from two angles, not two independent accounts.
  • Sources ⅚/36 (Twitter blog and SREcon talk) share authors; they corroborate wording, not independence.