Storage Tiering Service · View 08 of 31 · 3 · Structure
Decisions
- Serving and tiering are separate clusters, not separate namespaces. A bad tiering deploy, a noisy ClickHouse merge or a mover memory leak cannot evict a resolver pod.
- The resolver is a stateless Go service with an in-process bounded cache. It talks to vtgate and nothing else.
- Catalogue change data is streamed out through Vitess VStream into Kafka and ClickHouse. Classification reads that replica, so a nine-billion-row scan never runs against the store every read depends on (ADR-14).
Numbers
- 48 resolver pods across three rooms, sized for a 4× 60-second burst on cache hits alone. 64 Vitess shards, a primary and two semi-sync replicas each in DC-A, one async replica each in DC-B.
- 24 mover nodes with 2 × 25 GbE, for 900 TB a day (about 83 Gbps sustained) with headroom for rebuilds.
Assumption
- The resolver cache hits about 85% of the time at steady state, because sync clients and previews re-read the same objects. The proof phase measures it; the catalogue replicas are sized to survive a 50% hit rate.