Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
6 to work through
-
intermediate
A team must choose partition keys for an event stream. What does the choice determine, and what goes wrong when it is chosen carelessly?
2 min answer -
intermediate Multiple choice
After scaling consumers from 3 to 12, downstream data shows updates applied out of order. What happened?
2 min answer -
advanced
A location pipeline must choose a partition key. What are the candidates, and what does each optimise for?
2 min answer -
advanced
A stream requires ordering per user but not globally. How should partitioning be designed, and what breaks it?
2 min answer -
advanced
Alibaba reported a peak of 583,000 order creations per second during Singles' Day 2020. Take that as the sizing input for a log-based order-event pipeline with per-buyer ordering. Roughly how many partitions does the order topic need, and which assumption dominates the error?
2 min answer -
advanced
One customer generates 60% of your event volume. Their partition's consumer cannot keep up and the rest idle. What are your options?
2 min answer
3 terms in this topic
Key Skew
An uneven distribution of records across partitions, which caps throughput at the busiest partition regardless of how many exist.
conceptPartition Count Immutability
The partition count of a keyed log is effectively permanent, because raising it changes which partition a key hashes to and therefore breaks per-key …
conceptPartition Key Skew
One partition receiving disproportionate traffic because a small number of key values dominate, capping throughput at what a single consumer can handle.
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.