Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
5 to work through
-
beginner
A team runs 400 events per second through a pipeline with Kafka Streams for routing, Flink for windowed aggregation and a Spark job for a nightly correction pass, maintained by four engineers. Two of the three have had production incidents this quarter. What would you remove, what would you keep, and what would you leave alone even though it looks odd?
3 min answer -
intermediate Multiple choice
A team must join an order stream keyed by order id to a payment stream that carries order id but is keyed by payment id, with 12 partitions on one topic and 48 on the other, matching within seven days and with a standing requirement to reprocess the last 30 days after a logic fix. Which runtime fits and what is the deciding property?
3 min answer -
advanced
A Flink job checkpoints every 60 seconds. State has grown so that one checkpoint now takes 90 seconds to complete. No configuration changes. What happens over the next hour, and which metric will not show it?
3 min answer -
advanced
What justifies a stream processing framework over a simple consumer loop?
2 min answer -
advanced
What should drive the choice of stream processing framework for a high-volume recommendation platform?
2 min answer
2 terms in this topic
Checkpoint Interval
How often a stateful processor persists its state and offsets, which trades steady-state overhead against how much work is redone after a failure.
toolStream Processing Frameworks
Engines that process unbounded data with managed state, windowing and fault tolerance — chosen by state and time semantics rather than by throughput.
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.