intermediate 3 min answer Multiple choice

A creator analytics feature must show per-minute view counts for 8 million creators over a 90-day window at p95 under 500 ms while ingesting 300 thousand view events per second. Work out the row counts and decide which storage shape fits.

rollupretentioncolumnarsizingpinot
Pick one
Show the full answer Hide the answer

The assumptions, stated

Ingest is 300k events/s. There are 8M creators but only a fraction are receiving views in any given minute; assume 400k active creators per minute, which is the assumption to challenge first. A raw event compresses to roughly 40 bytes in a columnar store; a rollup row keyed by creator and minute with a handful of measures is roughly 35 bytes.

The arithmetic

  • Raw, per day: 300,000 × 86,400 ≈ 26 billion rows/day, about 1 TB/day compressed.
  • Raw, 90 days: ≈ 2.3 trillion rows, roughly 90 TB.
  • Rollup, per day: 400,000 active creators × 1,440 minutes ≈ 580 million rows/day, about 20 GB.
  • Rollup, 90 days: ≈ 52 billion rows, roughly 2 TB — a 45x reduction, and it is bounded by creators × minutes rather than by traffic, so it does not grow when a video goes viral.
  • Raw for 48 hours for drill-down: ≈ 52 billion rows, about 2 TB.

A per-minute rollup is also exactly the grain the feature displays, so no query has to aggregate anything it does not show.

Why the other options fail

  • Raw for 90 days aggregated at query time. A p95 of 500 ms over a 2.3-trillion-row scan at real concurrency is not a tuning problem; it is arithmetic. This is the right answer when the question set is unknown and volumes are far smaller — an internal exploration table at 3,000 events/s is 23 billion rows over 90 days and a warehouse handles it.
  • Per-hour rollups with no raw events. Cheapest and it deletes the product. Creators want to see the spike minute, and once raw data is gone, no reprocessing can recover per-minute detail. A pre-ingest rollup decision is irreversible in a way a retention change is not.
  • Raw for 90 days with an index on creator and minute. Row-store thinking. An index turns a scan into 2.3 trillion row lookups, and in a columnar store the equivalent is per-segment pruning, which helps the filter but still materialises every matching row before aggregating.

Which assumption dominates the error

The active-creator-per-minute count. If it is 2M rather than 400k, the rollup is 2.9 billion rows/day and 260 billion over 90 days — still viable, but now in the same order as the raw 48-hour table, and the retention split needs revisiting. The 40-byte row size, by contrast, is unlikely to be wrong by more than 2x, and the 90-day window is fixed by the product.

What the numbers rule in and out

They rule out any design that reads raw data on the user path, and they rule in a specialised store only because of concurrency and freshness rather than volume. Apache Pinot's documented requirement that rollup and dedup behaviour is fixed at table creation is the operational consequence: the rollup grain is chosen before the first event lands, and changing it later means a new table and a backfill.

When the specialised store is the wrong answer

If this feature were for 200 internal analysts rather than 8 million creators, a 5-minute materialised view in the existing warehouse answers it at a fraction of the operational cost. The specialised store is justified by concurrent user-facing queries against fresh data, not by the row count on its own.