pattern

Follower Fetching

also called Closest-Replica Fetch, Rack-Aware Consumption

Letting a consumer read from the in-sync replica in its own availability zone instead of from the partition leader, which removes consumer-side cross-zone transfer charges and adds replication delay to every read.

kafkakip-392cross-zonecostreplication

A streaming platform's cloud bill arrives and the largest line is not compute or storage but inter-zone data transfer. The arithmetic is unforgiving: a partition leader lives in one zone, and consumers are spread across three, so about two thirds of every byte consumed crosses a zone boundary and is billed twice, once leaving and once arriving.

At AWS list prices for 2025 to 2026, cross-AZ traffic inside a region costs about $0.01 per GB in each direction, so one GB moved costs roughly two cents. A topic serving 100 MB/s to consumers in other zones is on the order of $5,000 a month in transfer alone, before replication traffic is counted.

Follower fetching removes the consumer side of that. KIP-392, shipped in Kafka 2.4 (2019), extended the fetch protocol so a consumer may read from any in-sync replica rather than only the leader. The leader answers the first fetch with a preferred read replica in the consumer's own rack, and the consumer reads locally from then on.

Why it matters

This is one of the few streaming cost reductions that is a configuration change rather than an architecture change. For a read-heavy topic with several consumer groups, consumer fetch traffic is usually the largest single contributor to cross-zone charges, because it is the only component multiplied by the number of consumer groups.

It also reduces latency variance slightly, since the fetch no longer crosses a zone boundary, and it takes read load off leaders, which are the brokers already doing the most work.

Implementation patterns

  • Set replica.selector.class to RackAwareReplicaSelector on brokers and broker.rack to the availability zone id on each broker.
  • Set client.rack on every consumer to the zone id of the host it runs on, read from instance metadata at startup rather than from static configuration.
  • Verify, do not assume. Compare per-zone transfer metrics before and after, and check the preferred-read-replica the consumer actually selected. A consumer with no client.rack silently falls back to the leader and the saving quietly does not happen.
  • Pair it with rack-aware partition assignment (KIP-881, later versions) so consumer instances are assigned partitions that have a replica in their own zone in the first place.
  • Handle the producer side separately. Producers must write to the leader, so producer and replication traffic stay cross-zone. The remaining levers there are stronger compression, larger batches and single-zone producer placement for topics whose producers are co-located.

Industry example

The mechanism is documented in the Kafka project itself, and AWS's own guidance for Amazon MSK recommends setting client.rack to the availability zone id specifically to reduce network transfer cost and tail latency for consumer fleets. The managed-service recommendation is the clearest signal of how large the line item is in practice: it is unusual for a provider to publish instructions for spending less with it.

Failure scenarios

  • Reads sit behind the leader. A follower only serves data up to the high watermark it has learned from the leader, which propagates on fetch responses. Normally this is single-digit milliseconds. If a follower falls behind, the consumer reads older data while reporting healthy, and consumer lag measured against the leader's end offset grows for a reason that has nothing to do with the consumer.
  • A missing or wrong client.rack after an autoscaling group is rebuilt, a container image change or a new language SDK that does not implement the field. The cost returns and nothing alerts, because functionality is unaffected.
  • An out-of-sync replica is not eligible, so during a broker recovery consumers fall back to the leader and transfer cost spikes at the same time as everything else.
  • Zone failure interacts badly with placement. If consumers are pinned to a zone and that zone's replica is the one lost, the group fails over to leader reads and to cross-zone cost simultaneously.
  • Older client libraries ignore the preferred read replica entirely, which is a known gap in several non-Java implementations.

Trade-offs

Choose Gains Pays
Follower fetching on Consumer-side cross-zone transfer close to zero Reads lag the leader by replication and watermark propagation delay
Leader reads only Freshest possible read, one code path About two thirds of consumer bytes billed in both directions
Single-zone deployment No cross-zone traffic at all Loss of zone redundancy, which is usually unacceptable

When not to use it

When reads must reflect a write that just happened, read from the leader. Any read-your-writes requirement on the log, such as a consumer verifying its own produce before acknowledging a user, is not safe behind a follower whose watermark propagation is unbounded during a slow replica.

When the topic is low volume, skip it. A topic moving 1 MB/s costs a few dollars a month in cross-zone transfer, and the operational cost of rack labelling that has to stay correct through every deployment exceeds the saving. Apply this to the handful of topics that dominate the bill, and measure the bill first, because the intuition about which topic that is wrong more often than not.

Interview question

Q: Your Kafka bill shows $40,000 a month in cross-AZ transfer. You enable rack-aware follower fetching and the figure falls to $26,000. Where did the remaining $26,000 go, and what would you do next?

What a strong answer covers: producers still write to the leader wherever it is, so roughly two thirds of produce traffic is cross-zone, and replication to two other zones is cross-zone by definition and unavoidable at replication factor 3 with zone spread. Next steps in order of payoff: raise compression and batch size, which reduces every one of those streams at once; check whether any consumer is missing client.rack and therefore still reading from leaders; consider whether the replication factor and min.insync.replicas combination is actually required for every topic, since the telemetry topics may not need what the ledger topics need. A strong answer also names the thing not to do, which is collapsing to a single zone to save transfer, and states what it costs in availability.

Quick check

Quiz: After enabling follower fetching, which cross-zone traffic remains? Producer writes to the leader and all replication traffic, since only consumer fetches can be served locally.

Flashcard: What does a consumer give up by reading from an in-sync follower instead of the leader? Freshness: the follower serves only up to the high watermark it has learned from the leader, so reads sit behind by the replication and watermark propagation delay.