Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
4 to work through
-
beginner
A service keeps an in-memory map of product prices built by consuming a compacted topic, instead of calling the pricing service on each request. A colleague asks why the map is not just a cache. What is the honest answer, and what does the service now own?
2 min answer -
beginner
A team is told that a table and a stream are two views of the same thing. They have a Postgres table holding 12 million account balances and are asked to produce the stream of changes that produced it. Why does the equivalence only run one way in practice, and what has to exist before it runs the other way?
3 min answer -
intermediate
Review this design. Each of 18 services embeds a Kafka Streams GlobalKTable of a 9-million-row customer topic so that lookups are local, each instance holding it in RocksDB on local disk. Instances take 11 minutes to become ready. What would you remove, what would you change, and what would you leave alone even though it looks odd?
3 min answer -
advanced
What does stream-table duality mean practically, and how does it change system design?
2 min answer
3 terms in this topic
Changelog Stream
A stream whose records are keyed updates, so replaying it from the beginning reconstructs a table — the same information in the other of its two forms.
conceptCompaction Lag
The delay between a key being overwritten and the superseded record actually being removed, which decides how much larger a compacted topic is than i…
conceptStream-Table Duality
The equivalence between a stream of changes and a table of current state, where each can be derived from the other.
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.