Storage Tiering Service  ·  View 08 of 31  ·  3 · Structure

Container Architecture

Two Kubernetes deployments with different availability targets, different change cadences and different on-call expectations.

Editable source SVG draw.io All views
Serving plane · Kubernetes · 99.99% Object access Placement resolver Go · 48 pods Read router Go library Placement API register · delete Catalogue vtgate Vitess router MySQL shards 64 × 3 · semi-sync Release gate Go · delete role Recall Recall API Go · job handles Recall workflows Temporal Unpack workers Go · staging Tiering plane · Kubernetes · 99.5% Observe Kafka events · decisions ClickHouse aggregates · replica Catalogue CDC VStream → Kafka Move Mover workers Go · 24 nodes Pack builder Go · compactor Reconciler Go Decide Classifier jobs SQL + Go Ring controller Temporal Policy service PostgreSQL Ceph RGW tiers hot · warm · cold EOS + CTA tape archive Records management hold events point read replica copy Container Architecture Interface / broker Application we own Data store Security / platform Queue / topic External / third party synchronous event / async The mover's CAS commit to vtgate and the recall path to tape are drawn on views 18 and 20. v 1.0 · owner Storage Platform Architecture · date 2026-09

Decisions

  • Serving and tiering are separate clusters, not separate namespaces. A bad tiering deploy, a noisy ClickHouse merge or a mover memory leak cannot evict a resolver pod.
  • The resolver is a stateless Go service with an in-process bounded cache. It talks to vtgate and nothing else.
  • Catalogue change data is streamed out through Vitess VStream into Kafka and ClickHouse. Classification reads that replica, so a nine-billion-row scan never runs against the store every read depends on (ADR-14).

Numbers

  • 48 resolver pods across three rooms, sized for a 4× 60-second burst on cache hits alone. 64 Vitess shards, a primary and two semi-sync replicas each in DC-A, one async replica each in DC-B.
  • 24 mover nodes with 2 × 25 GbE, for 900 TB a day (about 83 Gbps sustained) with headroom for rebuilds.

Assumption

  • The resolver cache hits about 85% of the time at steady state, because sync clients and previews re-read the same objects. The proof phase measures it; the catalogue replicas are sized to survive a 50% hit rate.