Prefetch Skew
also called Unacknowledged Buffer Imbalance, Consumer Message Hoarding
The imbalance that appears when each consumer buffers a batch of messages locally - work queued behind a busy consumer cannot be taken by an idle one, so a pool stops self-balancing.
A rendering pool runs 12 workers against one queue with a prefetch of 100. Queue depth sits at 4,000, three workers are pinned at 100% CPU, and nine are idle. Throughput is roughly a quarter of what the pool can do.
Adding workers does nothing, which is the confusing part, because competing consumers is the pattern that is supposed to balance itself. It balances only while messages are unassigned. With a prefetch of 100, the broker has already handed 1,200 messages out, and the ones behind a heavy job are the property of a worker that will not look at them for hours. Jobs here range from 50 ms to 90 seconds, so one unlucky buffer can hold 2.5 hours of work while nine workers have none.
Why it matters
Prefetch exists to hide the broker round trip, and its correct value is a latency-bandwidth product, not a throughput knob: roughly the round-trip time divided by the per-message processing time, plus one. At 1 ms round trip and 200 ms jobs, the right prefetch is 1 or 2. At 1 ms round trip and 1 ms jobs, you need hundreds to keep the pipe full. Defaults of 100 or more exist because brokers are tuned for the second case, and most business workloads are the first.
The failure looks like under-capacity, so the usual response is to scale out. Scaling out makes it worse: more workers multiplied by the same prefetch means more messages hoarded, and the idle fraction grows.
Implementation patterns
- Set prefetch from the ratio, not the default. Long or variable tasks: 1 to 2. Short uniform tasks: hundreds.
- Separate queues by cost class so a 90-second job cannot be buffered in front of a 50 ms one. This fixes the variance that prefetch tuning can only soften.
- Keep the buffer inside the visibility timeout. If the time to drain a worker's buffer can exceed the timeout or the maximum poll interval, the broker reclaims the messages, redelivers them elsewhere and may trigger a rebalance — work gets done twice and progress stalls.
- In partitioned systems the equivalent lever is assignment, not consumer count. A slow key's partition cannot be helped by an idle consumer, so the knobs are records per poll and partition count.
- Alert on idle workers with non-empty queue depth. That single conjunction is the signature, and neither metric alone shows it.
Industry example
Any queue-backed media or document pipeline in production meets this within its first year of mixed workloads: thumbnailing at tens of milliseconds and full-length transcoding at minutes, sharing one queue and one prefetch value because they share a worker image. The fix that holds is two queues and two prefetch settings rather than a cleverer single number, because no single value is right for a bimodal distribution.
Failure scenarios
- Idle consumers alongside a deep queue, usually mistaken for a scaling problem.
- Redelivery storms when buffered messages outlive the visibility timeout, producing duplicate work and, in consumer-group systems, repeated rebalances.
- Latency inversion, where a small urgent job waits behind a buffer of large ones on one worker while capacity sits unused.
- Autoscaling oscillation, because queue depth stays high while CPU stays low, so the scaler adds workers that immediately hoard more messages.
- Poison messages hidden in a buffer, retried by one worker repeatedly and invisible in pool-level metrics.
Trade-offs
A low prefetch costs one broker round trip per message. At 1 ms that is a 0.5% overhead on a 200 ms task and a 50% overhead on a 2 ms task, which is exactly why the setting cannot be a house standard. Cost-class queues cost a second queue, a second deployment and a routing decision at publish time — real work, bought in exchange for a pool that balances without tuning.
When not to use it
When tasks are short, uniform and high-rate, a large prefetch is correct and skew is immaterial: every buffer drains in milliseconds, so an idle worker is never idle for long. It flips when the ratio between the longest and shortest task exceeds roughly an order of magnitude, which is the point where one buffer's contents stop being interchangeable with another's.
Interview question
Q: You have 20 workers on one queue, depth of 10,000, and average CPU across the pool is 15%. Nobody has changed the code. Where do you look, and what do you change first?
What a strong answer covers: the conjunction of idle workers and deep queue pointing at assigned-but-unworked messages rather than at capacity · prefetch and unacknowledged-message counts per consumer as the confirming evidence · the task duration distribution, because the real problem is variance · reducing prefetch as the immediate fix and splitting queues by cost class as the structural one · and the warning that scaling out first amplifies the symptom.
Quick check
Quiz: Twelve workers, prefetch 100, 4,000 messages queued, nine workers idle. What is wrong? The broker has already assigned the backlog; work queued behind long-running jobs cannot be stolen by idle consumers, so the pool no longer self-balances.
Flashcard: How should prefetch be chosen, and what is the symptom of getting it wrong on variable-length tasks? — Round-trip time divided by per-message processing time plus one, so 1 to 2 for long tasks and hundreds for very short ones; the symptom is idle consumers while the queue is deep, which scaling out makes worse.