Backup and Restore Service  ·  View 17 of 26  ·  5 · Runtime

In-Place Restore — Authorised, Snapshotted, Reversible

The most destructive thing the platform can do, and the steps that make it approved, recorded and reversible.

Editable source SVG draw.io All views
On-call SRE Owner approver Platform API Evidence store Restore orchestrator Target volumes Target database 1. in-place · point · reason 2. approval request 3. approve · WebAuthn, not the requester 4. intent recorded before action 5. start in-place workflow 6. fence writers · stop instance 7. snapshot all volumes · keep 7 d 8. restore over original 9. asserts pass · promoted 10. outcome · duration · snapshot id 11. done · revert available 7 d 12. revert → roll back snapshot In-Place Restore — Authorised, Snapshotted, Reversible If the snapshot fails, the workflow stops before the restore. There is no override flag. v 1.0 · owner Backup Platform · date 2026-09

Decisions

  • The approver must belong to the datastore's owner group and must not be the requester. Keycloak enforces WebAuthn for the approval, so the second person is a real second person.
  • The audit intent is written before the workflow starts. If the evidence write fails, the restore does not start. A destructive action with no record is worse than a delayed one.
  • The pre-restore Ceph RBD snapshot of every target volume is mandatory. If the snapshot fails, the workflow stops, and there is no flag to skip it.

Numbers

  • The target snapshot is kept for 7 days. Reverting takes minutes, because rolling back an RBD snapshot does not re-copy data.
  • In-place restores are expected to be rare, a few a year. Side-by-side restores followed by a service cutover cover most recoveries.

Risks

  • Datastores not on Ceph RBD (a few bare-metal hosts with local NVMe) cannot be snapshotted this way. They take a pgBackRest full copy of the target first, which is slower, and their stated in-place RTO includes it.