Storage Tiering Service  ·  View 22 of 31  ·  6 · Operations

Deployment

Two data centres with three fault-isolated rooms in the primary, a witness for quorum, and storage that replicates independently of the service.

Editable source SVG draw.io All views
DC-A · primary · three rooms Serving and tiering Kubernetes clusters serving · tiering Vitess primaries 64 shards · semi-sync Kafka · ClickHouse 6 + 12 nodes Storage tiers Ceph hot · warm · cold RGW multisite zone A Tape library A EOS buffer · CTA DC-B · secondary Standby Kubernetes serving warm · resolvers up Vitess replicas async · RPO ≤ 10 s Storage tiers Ceph hot · warm · cold RGW multisite zone B Tape library B second copy Witness site etcd voter Vitess topology Healthchecks dead-man monitor Offsite vault Phase 3 · cartridges binlog multisite sync Deployment — Two Data Centres and a Witness Security / platform Data store Queue / topic External / third party event / async In-DC failover of a shard primary is VTOrc, under 60 s. Moving serving to DC-B is a declared runbook with a 15 min RTO. v 1.0 · owner Infrastructure · date 2026-09

Decisions

  • Rooms in DC-A play the role of zones: shard primaries and their two semi-sync replicas are in different rooms, so a room loss is an automatic VTOrc failover rather than a site event.
  • DC-B runs resolvers warm against async replicas and a second Ceph zone. Promoting DC-B is a rehearsed runbook, not an automatic action, because a split-brain catalogue is worse than 15 minutes of failed reads.
  • The witness site holds the third etcd voter for Vitess topology and the dead-man monitor, so neither depends on the data centre it is watching.

Numbers

  • Logical bytes by tier at the planning mix: about 12 PB hot, 8 PB warm, 12 PB cold, 8 PB archive. Tier sizing and hardware belong to the storage team; the service consumes them.
  • In-DC catalogue RPO 0, RTO 60 s. Cross-site RPO 10 s, RTO 15 min.

Risk

  • Ceph multisite replication is asynchronous. The release gate's check for the destination replica in DC-B is what makes a site loss unable to delete the last copy, so that check is tested in the proof phase by killing DC-A mid-release.