Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.
5 to work through
-
intermediate Multiple choice
A creator analytics feature must show per-minute view counts for 8 million creators over a 90-day window at p95 under 500 ms while ingesting 300 thousand view events per second. Work out the row counts and decide which storage shape fits.
3 min answer -
intermediate Multiple choice
A team proposes a specialised real-time analytics store for an internal dashboard used by twelve analysts. Assess.
2 min answer -
advanced
A platform needs sub-second analytical queries over a continuously updating dataset. What kind of store, and what are the trade-offs?
2 min answer -
advanced
A platform needs sub-second analytical queries over recent high-volume event data. What store characteristics matter?
2 min answer -
advanced Multiple choice
A telemetry product must answer arbitrary filtered queries over hundreds of billions of events retained for 15 months, with sub-10-second response for the last 24 hours. Each event has 40 to 200 fields and cardinality is unbounded. Which storage design fits?
3 min answer
5 terms in this topic
Primary-Key-Partitioned Upsert
Routing every version of a record to one partition so a real-time store can resolve the latest value locally - which buys sub-second latest-value que…
toolReal-Time Analytics Store
A database built for sub-second aggregate queries over freshly ingested data at high concurrency, serving user-facing analytics rather than internal …
practiceRollup Policy
The decision - made before ingestion, not after - about what granularity of data is retained for how long, and at what point raw events are collapsed…
toolSub-Second Aggregation Store
A database built to ingest continuously and answer aggregate queries over recent data in milliseconds, occupying the gap between OLTP and the warehouse.
patternUnbundled Event Store
A telemetry store that separates writing, compaction and querying into independently scaled services over shared object storage and a metadata store,…
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.