pattern

Zone-Aware Routing

also called Topology-Aware Routing, Same-Zone Preference, Locality-Aware Load Balancing

Sending a request to a replica in the caller's own availability zone and crossing a zone boundary only when the local replicas cannot serve - which removes a round trip and a per-gigabyte charge from every internal call at the price of uneven load and a sharper failover edge.

availability-zonesload-balancinglatencycross-zone-costfailover

A service mesh that spreads calls evenly across three zones is doing something nobody asked for. Every internal hop pays a cross-zone round trip of roughly 1 to 2 ms instead of a same-zone hop of 0.1 to 0.5 ms, and on AWS it pays $0.01 per GB in each direction (list price, 2026) which is the one network charge where sender and receiver are both billed. A request that makes 30 internal calls pays that round trip 30 times, and two thirds of those crossings were avoidable.

Zone-aware routing makes the caller prefer targets in its own zone, spilling over to other zones only when local healthy capacity runs short. The idea appears under different names in different layers: Envoy's zone-aware routing, Kubernetes topology-aware hints, Kafka's rack-aware follower fetching, and the cross-zone toggle on a Network Load Balancer.

Why it matters

The two savings are independent and both large at scale. Latency compounds through a call graph: removing 1 ms from each of 30 sequential hops removes 30 ms from p50 and more from the tail, because each crossing carries its own tail. Cost scales with volume rather than request count: a service exchanging 250 TB a month across three zones crosses a boundary on about two thirds of it, which at $0.02 per GB round-trip is roughly $3,000 a month for a benefit nobody can name.

The third reason matters most. A same-zone call path means a zonal failure removes one vertical slice rather than degrading every slice, which is the property cell-based designs exist to get. Uniform cross-zone meshing guarantees the opposite: every request touches every zone, so one bad zone touches every request.

Implementation patterns

  • Spill-over rather than strict affinity. Prefer local targets, then send remote once local healthy capacity falls below a threshold. Envoy expresses this as a percentage of expected local capacity; the strict version turns a partial zone problem into a hard failure.
  • Let the platform do it. Kubernetes topology hints and the NLB cross-zone setting are cheaper than application logic. An Application Load Balancer always balances across zones at the balancer level and that hop is not charged; a Network Load Balancer disables cross-zone by default precisely because enabling it creates cross-zone charges.
  • Pin storage reads locally. Kafka's rack-aware fetching (KIP-392) lets a consumer read an in-sync replica in its own zone instead of the leader.
  • Keep the NAT path zonal. One NAT gateway serving three zones adds a cross-zone hop on top of the $0.045 per GB processing charge (us-east-1 list, 2026); one gateway per zone removes the hop and a VPC endpoint removes the processing charge for the services it covers.

Industry example

Kafka's rack-aware fetching exists because the cross-zone bill on a high-throughput cluster is comparable to the compute bill. With client.rack set on consumers and a rack-aware replica selector on the brokers, the leader answers the first fetch with a preferred read replica in the consumer's own zone and consumer-side cross-zone transfer falls to near zero. Producers still write to the leader wherever it lives, so roughly a third of produce traffic and all replication traffic remain cross-zone — locality optimises one direction of a system at a time.

Failure scenarios

  • Strict affinity with uneven capacity. Two instances in zone A and four in zone B gives A's callers half the capacity of B's. The symptom is per-zone latency divergence with no unhealthy targets, and the aggregate dashboard shows nothing.
  • Locality during a zonal degradation. Local replicas are up and slow, so affinity keeps feeding them and the caller never finds the healthy capacity one zone away. This needs latency-based outlier detection, not health checks.
  • Spill-over threshold set too high, so a small local dip floods other zones at once and the cost spike arrives as a bill rather than an alert.
  • Stateful reads against a local follower, so a read-your-writes path silently returns stale data to save a round trip.

Trade-offs

Choose Gains Pays
Zone-aware with spill-over Lower p50 and p99 on chatty paths; most cross-zone charges removed; vertical failure slices Capacity must be balanced per zone; a second routing mode to reason about during an incident
Uniform cross-zone Even load; simplest mental model; no per-zone capacity planning A round trip and a per-GB charge on every internal call; every request exposed to every zone

When not to use it

Skip it when the traffic is neither chatty nor large. A service at 200 requests per second with one downstream call moves a few gigabytes a month; locality saves a few dollars and costs a configuration surface that will confuse somebody mid-incident. Skip it also when you cannot hold balanced capacity in every zone, because a fleet of five instances across three zones cannot be locality-routed without creating a hot zone.

Above all, do not apply it to the balancer fronting external traffic. Cross-zone balancing there is what makes a zone's instance count irrelevant to its share of traffic, and on an ALB it is free. Locality belongs on internal service-to-service and client-to-storage paths.

Interview question

Q: Your mesh spreads every internal call evenly across three zones and finance asks why network charges are 15% of the cloud bill. What would you change, what would you measure first, and what would you deliberately leave crossing zones?

What a strong answer covers: measure per-service cross-zone bytes rather than the total, because the charge is almost always concentrated in two or three chatty paths. Then zone-aware routing with spill-over on those paths, zonal NAT gateways, VPC endpoints, and rack-aware consumer fetching. Deliberately leave crossing: the external balancer, the database's synchronous replication, and any path where the local replica is not authoritative. State the precondition — locality needs balanced per-zone capacity and headroom to absorb spill-over, or the saving is paid back in tail latency.

Quick check

Quiz: A service makes 30 internal calls per request and the mesh routes uniformly across three zones. Roughly how much latency does locality remove, and which hop should still cross? Answer: about 1 to 2 ms on each of roughly 20 crossings, so 20 to 40 ms off p50 and more off the tail; the database's synchronous replication hop should still cross, because that crossing is the durability guarantee.

Flashcard: Why is cross-zone traffic the only cloud network charge where both sides pay? — The provider bills $0.01 per GB leaving a zone and $0.01 per GB entering another (AWS list price, 2026), so one gigabyte between zones costs $0.02 and lands on both services' cost attribution.