Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
3 to work through
-
intermediate
A job enriches order events with customer tier by looking up a table built from a compacted topic. After a routine deploy, 3% of orders in the first ten minutes come out with tier null, then the rate returns to zero. Throughput and error rate are flat. What is happening?
2 min answer -
advanced
A job interval-joins click events to impression events with a two-hour window. To fix a parser bug the team resets only the click consumer's offsets to three days ago and lets the running job catch up. Walk through what happens, what the output looks like, and what stops it.
3 min answer -
advanced
A pipeline must join two streams whose events arrive at different times. What makes streaming joins hard, and what are the options?
2 min answer
5 terms in this topic
Join Replay Misalignment
The failure that follows from replaying one input of a stateful join while the other stays live - event times no longer overlap, window state is neve…
patternStream Enrichment
Attaching reference data to a stream by lookup against a materialised table rather than by joining two unbounded streams.
conceptStreaming Join
Combining two streams, or a stream and a table, where the records to be joined may not arrive at the same time or in the same order.
conceptTable Bootstrap Lag
The interval after a stream job starts during which its lookup table is only partly loaded, so enrichment lookups miss for keys that appear late in t…
patternTemporal Join
A join that enriches an event with the reference value as it was when the event occurred, rather than as it is now - which is the difference between …
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.