Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
5 to work through
-
beginner Multiple choice
An hourly batch job writes each hour's aggregate into a dt/hour partition and is safe to rerun after any crash. The team reimplements the same logic as a streaming job writing to the same table. The streaming job is at-least-once and now produces duplicate rows after every restart. What made the batch version's retries safe?
2 min answer -
intermediate
A product manager asks for a "real-time" dashboard. How do you turn that into a design decision?
2 min answer -
intermediate
A quick-commerce platform must decide whether a data flow should be streaming or batch. What actually decides it?
2 min answer -
intermediate
Which parts of a grocery platform's data processing genuinely require streaming, and which are better as batch?
1 min answer -
advanced
A business sponsor asks for a real-time data platform because "the competition has one". Reporting is currently a nightly batch that lands at 06:00 and nobody has complained. How do you handle this?
2 min answer
3 terms in this topic
Freshness Requirement
How stale data may be before the decision it supports degrades — the only question that justifies streaming over batch.
patternMicro-Batching
Executing a stream as a sequence of small bounded jobs on a fixed trigger interval, which buys batch failure semantics and atomic sink commits at the…
practiceStreaming vs Batch Decision
Choosing between continuous and periodic processing based on the decision latency the business actually requires, not on the appeal of real-time.
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.