Compaction Lag
also called Cleaner Lag, Dirty Segment Lag
The delay between a key being overwritten and the superseded record actually being removed, which decides how much larger a compacted topic is than its key set and how long a rebuild from it takes.
A service is expected to restore a local table of 9 million customers in about a minute. It takes eleven. Nobody has done anything wrong: the topic is compacted, the key count is right, and the record size estimate is right. The missing term is when compaction actually happens.
Compaction is not a property that holds continuously. It is a background process with thresholds, and the interval between a write superseding an older value and that older value being physically removed is compaction lag. A compacted topic is a table plus however much duplicate history the cleaner has not reached yet.
Why it matters
Restore time is a function of records read, not of distinct keys. Every duplicate still in the log is read and applied during a rebuild, so compaction lag is directly a rebalance-time and deploy-time cost — paid on every deployment, and again during the incident when instances are being replaced.
Kafka's defaults make the gap substantial. The active segment is never compacted, and a closed
segment is only cleaned once the ratio of uncleaned to total log exceeds
min.cleanable.dirty.ratio, which defaults to 0.5. A topic can therefore legitimately hold about
twice as many records as it has distinct keys, indefinitely, while reporting as compacted.
Implementation patterns
- Size rebuild time from the topic's byte count, never from keys times record size. The byte count is the only number that reflects the lag.
- Tune the trade deliberately. Lowering
min.cleanable.dirty.ratiocompacts more eagerly and shortens rebuilds at the cost of continuous cleaner I/O;segment.msandsegment.bytesgovern how quickly segments close and become eligible at all. - Use
max.compaction.lag.mswhere a bound matters for correctness or regulation, so a key's superseded values cannot linger indefinitely in a low-traffic topic. - Bootstrap from a snapshot instead of the log where restore time is on the critical path, and keep warm standby replicas so a rebalance promotes rather than replays.
- Remember tombstones have their own clock.
delete.retention.msdefaults to 24 hours, so a consumer resuming from a stored offset after a longer outage can miss a deletion entirely and keep a row the rest of the world removed.
Industry example
The mechanism is documented in Apache Kafka's own design notes on log compaction, and the defaults above are the shipped values rather than folklore — which means every cluster running in production has them unless someone changed them deliberately. It surfaces most often in the pattern where many services each materialise a shared reference topic into a local store: the topic looks small, the restore takes minutes, and the fleet's readiness time is dominated by a broker setting that nobody in the consuming teams has read.
Failure scenarios
- Deploys that take an hour because every instance replays twice the data anyone budgeted for.
- Disk pressure on brokers for a topic believed to be bounded by its key count.
- A rebuild that resurrects deleted keys, when the rebuild starts after the tombstones have aged out.
- Nondeterministic tables if the partition count was ever changed, since compaction is per-partition and an old value can survive in the partition the key no longer maps to.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Eager compaction (low dirty ratio) | Small topic and fast rebuilds | Continuous cleaner CPU and disk I/O on brokers |
| Default 0.5 | Cheap brokers | Up to about 2x the records and rebuild time |
| Snapshot-based bootstrap | Restore decoupled from the log entirely | A snapshot mechanism to build and keep correct |
When not to use it
The concept only earns attention when something rebuilds from the topic. If the compacted topic exists purely so that a late-joining consumer can eventually catch up, and no service's readiness depends on finishing that read, the lag is a storage detail and tuning it is wasted effort. And if restore time is genuinely on the critical path, the better answer is usually to stop restoring from the log at all — a snapshot plus standby replicas beats any cleaner setting.
Interview question
Q: A service materialises a 9-million-key compacted topic into a local store and takes eleven minutes to become ready, against an estimate of one. Where does the time go, and what would you change in what order?
What a strong answer covers: restore reads records rather than keys, and the active segment plus the 0.5 dirty-ratio default mean far more records than keys · measure the topic's actual bytes before theorising · the ordered fixes: warm standbys so a rebalance does not restore at all, then snapshot bootstrap, then broker-side compaction tuning last because it affects every consumer · and the tombstone-retention trap for consumers that resume after a long outage.
Quick check
Quiz: Why can a compacted topic hold roughly twice as many records as it has distinct keys?
The active segment is never compacted and closed segments are only cleaned once the dirty ratio
exceeds min.cleanable.dirty.ratio, which defaults to 0.5.
Flashcard: Your compacted topic has 9 million keys. What number should you use to estimate rebuild time? — The topic's actual byte count, because compaction lag means the log still holds uncompacted duplicates and a rebuild reads every one of them.