beginner 2 min answer

Six backends sit behind one load balancer: two in zone A and four in zone B. The zone A pair runs at 70% CPU while the zone B four sit at 25%. The load balancer's per-zone request counts are even and no target is unhealthy. What is happening?

load-balancingavailability-zonescross-zonecapacityskew
Show the full answer Hide the answer

The diagnosis

With cross-zone load balancing off, a load balancer node sends traffic only to targets in its own zone. Clients are spread across the zones' entry addresses roughly evenly, so traffic splits per zone, and each zone's share is then divided among however many targets that zone happens to hold.

The arithmetic

Zone A takes half the traffic and splits it two ways: 25% per instance. Zone B takes half and splits it four ways: 12.5% per instance. The per-instance ratio is 2:1, which is exactly what the CPU numbers show. Nothing is broken; the fleet is simply asymmetric and the balancer is not compensating.

The misleading signal

The even per-zone request count is the clue people discard. Even per zone is the cause, not the refutation. A dashboard that charts requests by zone looks correct while a dashboard that charts requests per target shows the skew immediately, which is why per-target rate belongs on the ingress dashboard rather than per-zone totals.

The fix and what it costs

  • Keep target counts equal per zone. Make the autoscaler add capacity per zone instead of to the fleet as a whole. This is free and it is the common fix.
  • Or turn cross-zone balancing on, which evens per-target load at the cost of inter-zone data transfer charged by the gigabyte in each direction, plus a small latency add on the cross-zone hop.

Decision rule: keep cross-zone off and the zones symmetric when traffic volume makes transfer charges material; turn it on when symmetry cannot be maintained - small fleets, mixed instance sizes, or capacity that disappears unpredictably.

The failure mode to name is the one that follows: during a peak the hot zone saturates first, somebody scales the fleet, and most of the new capacity lands in the zone that was already idle.

When this is the wrong answer

If cross-zone balancing is already on, this explanation does not apply and the next suspects are long-lived connections pinned at connect time (one connection carrying very different request volumes), a hash-based algorithm with a skewed key, or a backend that is slower rather than busier. The general habit: before blaming the algorithm, check whether the pool the algorithm sees is the pool you think it sees.