metric

Consumer Lag

The gap between the latest offset written to a partition and the offset a consumer group has processed, expressed in records or in time.

streamingmonitoringoperations

The single most important operational signal in a streaming system, and the one most often monitored badly.

Measure it in time, not records. A lag of 50,000 records means nothing without a rate: it is two seconds on one topic and nine hours on another. Time-based lag is directly comparable to a freshness objective, which is what consumers actually care about.

The alerting subtlety is that absolute lag is a poor trigger on its own. A brief spike after a deployment or a rebalance is normal; a small but steadily growing lag is the genuine emergency, because it means processing rate is below arrival rate and the system will not recover on its own. Alert on sustained growth in lag over a window, and separately on lag exceeding the freshness objective.

Two things that look like healthy lag and are not. Zero lag on a partition receiving no data may mean the producer has stopped, which requires a separate throughput check. And lag that resets to zero without processing usually means offsets were reset or the retention window expired and the consumer silently skipped ahead — data loss that appears on a dashboard as a return to health.