Backup and Restore Service  ·  View 23 of 26  ·  6 · Operations

Roadmap — Proof First, Breadth Second

What gets built in which order, and why the first release covers fewer engines but proves recovery for all of them.

Editable source SVG draw.io All views
MVP · 2 quarters PostgreSQL + volumes Tier 1 and 2 Custody A + B locked Daily Tier 1 rehearsal depth 3 Manual restore drill Phase 2 · 2 quarters MySQL · TiDB · ClickHouse Tier 3 sampling Depth 4 for Tier 1 Self-service restore Phase 3 Tape custodian rehearsed Clean-point search Provable disposal Restore capacity reserve Roadmap — Proof First, Breadth Second Application we own Data store Opportunity Security / platform Nothing is called protected until the MVP's rehearsal has run for 30 consecutive days. v 1.0 · owner Platform Architecture · date 2026-09

Decisions

  • The MVP ships the rehearsal loop with the first engine, not after the last. A platform that captures six engines and proves none is the problem this project exists to fix.
  • PostgreSQL and Kubernetes volumes come first because they are 44 of the 120 datastores and 5 of the 6 Tier 1 stores.
  • The tape custodian is Phase 3 in the requirement, but tape hardware has the longest lead time. The library is ordered in the MVP so that Phase 3 is configuration work, not procurement.

Exit criteria

  • MVP: 30 consecutive days with every Tier 1 datastore proven daily, one timed manual restore with the control plane switched off, and one catalogue rebuild from a scan, diffed clean.
  • Phase 2: every datastore enrolled, zero unprotected resources older than 72 hours, first quarterly gameday published.

Risks

  • Until Phase 2, MySQL, TiDB and ClickHouse keep their existing backups. The report marks them 'legacy, unproven' rather than leaving them off, so the gap stays visible.