Storage Tiering Service · View 14 of 31 · 4 · Data
Decisions
- The resolver emits the first read of each object by each reader class in each hour, with a count, instead of every read. Classification needs last access and distinct access days, not a range-read log (ADR-13).
- Objects under 1 MB are aggregated per folder cohort at ingest, which is the unit they will be classified and moved as.
- A canary event per Kafka partition every 60 seconds. The freshness gate and the external dead-man monitor both read its age, so the alarm and the brake can never disagree.
Numbers
- Estimated 60,000 to 90,000 events a second at peak after deduplication, about 120 billion a month at 180 bytes each: roughly 22 TB raw a month, 3 to 5 TB compressed in ClickHouse.
- Raw events kept 90 days, aggregates three years. Demotion stops when the newest canary is more than 15 minutes old.
Rejected
- Flink or Spark streaming for aggregation. ClickHouse materialised views over a Kafka engine table do the rollup at ingest, and one fewer distributed system is one fewer thing to be the dominant controllable cost.