Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
4 to work through
-
advanced
A finance-reporting platform runs a nightly batch job and a streaming job that compute the same daily revenue figure, and the batch number is the one the controller signs. You are asked to retire the batch path. Sequence it so every step is reversible and say where the point of no return is.
3 min answer -
advanced
A platform maintains separate batch and streaming implementations of the same logic. What does that cost, and what are the alternatives?
2 min answer -
advanced
A publisher keeps every piece of published content in one ordered Kafka log and rebuilds every downstream store by replaying it, as The New York Times described in 2017. What does that buy, what does the ordering requirement cost, and where would copying it be a mistake?
3 min answer -
advanced
Walk me through how you would decide whether a new analytics platform needs a batch processing path at all. The team already runs a stream processor, and nobody has asked for historical reprocessing yet.
3 min answer
4 terms in this topic
Denormalised Log
A second topic that republishes each record with its referenced data already resolved, so consumers read complete payloads instead of each performing…
patternKappa Architecture
Treating the stream as the single processing path, and handling historical reprocessing by replaying the log through the same code.
patternLambda Architecture
Running the same computation as both a scheduled recomputation over all retained history and a streaming job over recent events, so that any defect i…
patternReplay Pipeline
A single processing path that produces both live and historical results by re-running the same code over retained input, replacing the two-path Lambd…
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.