Enabling tiered storage let a platform raise log retention from 7 days to 90 days while shrinking broker disks and cutting the cluster bill roughly in half. Four months later the first 30-day reprocessing run takes 14 hours against the 3 hours the team projected from local-disk throughput. What did the saving actually buy and when does the bill arrive?
Show the full answer Hide the answer
What the saving bought
Retention stopped being a function of broker disk. At 200 MB/s ingest, 90 days at replication factor 3 is about 4.5 PB of provisioned block storage; in the remote tier the data is held once rather than three times, on storage that costs roughly a tenth to a twentieth per GB-month. The brokers can then be sized for throughput and page cache instead of capacity, which is where the rest of the halving comes from.
That is a real and large win, and it is a win on storing history.
What it pays
Replay throughput and first-byte latency, because a remote read shares nothing with a tail read.
- No page cache. A consumer at the head of the log is usually served from memory. A consumer 60 days back causes the broker to fetch whole segments from object storage before it can serve them, so per-fetch latency moves from sub-millisecond to tens or hundreds of milliseconds and throughput is set by the remote-fetch path rather than by disks.
- The path is deliberately rate-limited. Kafka's tiered storage quotas (KIP-956, shipped with
the 3.9 release in which tiered storage went GA) give brokers
remote.log.manager.fetch.max.bytes.per.secondas a global cap across every partition fetching from remote. The default is effectively unlimited, and any operator who has watched a replay evict the page cache for live consumers sets it. After that, replay speed is a configured number, not a hardware number. - Request and transfer charges. Object storage bills per GET and per GB moved; a 30-day replay reads the whole window, once per consumer group that replays it.
- The design assumption. Kafka's 3.9 operations documentation describes remote reads as infrequent - backfill or failure recovery. The feature was built to make long retention affordable, not to make long replays fast.
The number to publish
Retention is meaningless on its own. Publish retention next to measured replay throughput from the remote tier, on day one, before anyone needs it: "90 days retained, replayable at 400 MB/s, so a full replay is about 14 hours." Then keep local retention at least as long as the reprocessing you do routinely - if a weekly correction run covers 48 hours, hold 48 hours on broker disk. For genuine full-history recomputation, export the log once to a columnar table and reprocess there with warehouse parallelism, rather than pulling petabytes back through brokers built for streaming.
The bill arrives at the worst moment by construction: not at enablement, but months later, during an incident, when the estimate everyone quoted came from local-disk numbers.
When this is the wrong answer
For a topic nobody replays - compliance archives, audit logs kept because a regulator requires them - the replay cost is theoretical and tiered storage is simply cheaper. And if retention is already only days and the broker disks are small, adding a remote tier adds a remote log manager, a bucket, lifecycle policies and a new class of failure for no saving. Tiered storage earns its place when retention is long and reads of old data are rare.