A marketplace platform at Flipkart's scale runs three separate Kafka clusters for orders, device telemetry and experimentation. Consolidating them into one shared cluster cuts the infrastructure bill by about 40%. What did that 40% buy, and when does the bill arrive?
Show the full answer Hide the answer
What is gained, quantified
Three clusters each need a minimum viable shape: three brokers for replication factor 3, controller quorum, monitoring, and enough headroom that a single broker loss is survivable. The minimum is paid three times while the average utilisation of each cluster sits somewhere around 25%. One cluster running at 60% utilisation on a larger fleet does the same work with fewer brokers, one upgrade path, one set of runbooks and one on-call rotation that actually gets practice. The 40% is real and it recurs monthly.
What is paid
- Quotas stop being optional. Brokers share network threads, request handler threads and page cache. A telemetry backfill replaying 30 days at full speed will starve order-topic produce latency, because a fetch for old data evicts the page cache the order topic was being served from. Producer and consumer byte-rate quotas per client id become a correctness control, not a tuning nicety.
- Failures correlate. A rolling broker restart now triggers consumer-group rebalances in all three domains at the same moment. The orders team's p99 moves when the experimentation team deploys.
- One bad client reaches everyone. A producer in a retry loop with no backoff occupies request handler threads, and every tenant sees elevated produce latency with no error anywhere.
- Nobody owns the retention bill. Storage is retention × throughput × replication factor. At 200 MB/s and 30 days with RF3 that is roughly 1.5 PB, and in a shared cluster it appears as one line item that belongs to the platform team, so no product team ever shortens a retention.
When the bill arrives
At the first cross-domain incident, and at the first quarter where the bill grows faster than event volume. Both are predictable and neither shows up in the migration's success metrics. The incident is the expensive one: a shared cluster converts a single-team problem into an all-teams problem, and the post-incident instinct is to split back out under pressure, which is the worst time to do it.
How to keep the option to reverse
Do the consolidation in a way that makes splitting a cutover rather than a redesign. Prefix topic names by domain, issue separate client credentials per domain, set per-domain quotas on day one even when there is headroom, and keep the per-domain dashboards. Then a split is mirror-and-switch. Teams that merge into one flat namespace with one shared service account find that the split requires renaming every topic in every consumer.
When not to consolidate
Keep a workload separate when its availability target or its retention policy differs from the others, or when it sits on the path that takes money. Orders have a different failure cost from clickstream, and a shared cluster quietly gives every tenant the availability of the noisiest one. The honest version of this decision is two clusters, not one or three: the money path alone, and everything else together. That captures most of the 40% and leaves the blast radius where it belongs.