Backup and Restore Service · View 19 of 26 · 6 · Operations
Decisions
- Custody clusters have their own racks, switches, Ceph admin keys, bastions and Keycloak realm. They share no administrative trust with production or with the backup platform.
- The control plane in DC-B is warm: the Argo CD application is defined and the catalogue standby is streaming, but workloads are scaled to zero. Starting it is a runbook step that gamedays measure, not a claim.
- The isolated enclave sits in DC-B next to custody B. Rehearsals that read from A cross the data-centre link, which is how the cross-site restore throughput gets measured.
Numbers
- Custody A: 18 hosts × 24 × 22 TB HDD, EC 8+3, about 9.5 PB raw. Custody B: 14 hosts, about 7.4 PB raw. Both are planned at 75% fill.
- Backup network: 2 × 100 GbE per capture and gateway host, about 25 Gbit/s for the 3 GB/s capture target and 32 Gbit/s for the 4 GB/s restore target.
- Rehearsal cluster: 16 nodes × 4 × 7.68 TB NVMe, enough for all six Tier 1 stores restored at once.
Risks
- Both data centres are in one metro area. A regional disaster leaves the tape vault as the only surviving custodian, with a stated RTO measured in days. That is accepted for now and flagged for the board-level risk register.