Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
5 to work through
-
intermediate
A marketplace platform at Flipkart's scale runs three separate Kafka clusters for orders, device telemetry and experimentation. Consolidating them into one shared cluster cuts the infrastructure bill by about 40%. What did that 40% buy, and when does the bill arrive?
2 min answer -
intermediate Multiple choice
A topic ingests 200 MB per second, replication factor 3, retained 30 days. Producers and consumers are spread across 3 availability zones with no rack awareness in the consumers. Roughly what does the storage footprint come to?
3 min answer -
advanced
A connected-vehicle telemetry pipeline's cost is growing faster than the fleet. Where does the cost go, and what reduces it?
2 min answer -
advanced
A streaming pipeline's cost grows faster than its event volume. What are the drivers?
1 min answer -
advanced
A streaming platform's cloud bill is dominated by a line item nobody recognises: inter-zone data transfer. Explain and fix.
2 min answer
4 terms in this topic
Follower Fetching
Letting a consumer read from the in-sync replica in its own availability zone instead of from the partition leader, which removes consumer-side cross…
metricRetention Cost
The storage bill for keeping a log replayable, which is set by retention multiplied by throughput multiplied by the replication factor.
conceptStreaming Cost Model
The cost structure of an always-on streaming system, where compute runs continuously regardless of volume and cross-zone data movement is a major lin…
patternTiered Storage
Keeping recent log segments on broker local disk and older ones in object storage, which decouples retention from broker sizing and turns replay from…
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.