Backup and Restore Service  ·  View 19 of 26  ·  6 · Operations

Deployment — Two Data Centres and a Vault Site

Where every component runs, which failures each placement survives, and which network carries backup traffic.

Editable source SVG draw.io All views
Data centre A · on-premises bnr-control · active Control cluster API · Temporal Catalogue primary Patroni · sync replica Custody A · own racks, own admins RGW gateways 8 · behind HAProxy OSD hosts 18 × 24 HDD · EC 8+3 Data centre B · on-premises bnr-control · warm Control cluster scaled to zero Catalogue standby Patroni standby Custody B · own racks, own admins RGW gateways 6 OSD hosts 14 × 24 HDD · EC 8+3 Isolated enclave · no route out Rehearsal cluster 16 NVMe nodes Vault site Tape library LTO-9 WORM OpenBao voter 5th of 5 Dead-man monitor shared observability Paging Alertmanager streaming custody copier Deployment — Two Data Centres and a Vault Site Security / platform Data store Interface / broker event / async Backup traffic rides a dedicated 2 × 100 GbE network, never the production service network. v 1.0 · owner Infrastructure · date 2026-09

Decisions

  • Custody clusters have their own racks, switches, Ceph admin keys, bastions and Keycloak realm. They share no administrative trust with production or with the backup platform.
  • The control plane in DC-B is warm: the Argo CD application is defined and the catalogue standby is streaming, but workloads are scaled to zero. Starting it is a runbook step that gamedays measure, not a claim.
  • The isolated enclave sits in DC-B next to custody B. Rehearsals that read from A cross the data-centre link, which is how the cross-site restore throughput gets measured.

Numbers

  • Custody A: 18 hosts × 24 × 22 TB HDD, EC 8+3, about 9.5 PB raw. Custody B: 14 hosts, about 7.4 PB raw. Both are planned at 75% fill.
  • Backup network: 2 × 100 GbE per capture and gateway host, about 25 Gbit/s for the 3 GB/s capture target and 32 Gbit/s for the 4 GB/s restore target.
  • Rehearsal cluster: 16 nodes × 4 × 7.68 TB NVMe, enough for all six Tier 1 stores restored at once.

Risks

  • Both data centres are in one metro area. A regional disaster leaves the tape vault as the only surviving custodian, with a stated RTO measured in days. That is accepted for now and flagged for the board-level risk register.