Backup and Restore Service  ·  View 26 of 26  ·  7 · Assurance

Failure Modes — Detection, Containment, What Keeps the Copies Safe

Every named failure class from the requirement, how it is noticed, what limits it, and the mechanism that means the copies survive it.

Editable source SVG draw.io All views
Detect Contain Protected by Recover Version drift Engine version seen Capture refused Certified pairs Certify, then resume Verifier silence No heartbeat Page at 2× interval Monitor outside Rerun missed proofs Catalogue loss Patroni failover fails Restore by envelope Self-describing artefacts Rebuild by scan Old logical corruption Invariant fails Stop in-place restores 35-day PITR · 12 monthly Clean-point search Stolen capture identity Denied deletes logged Revoke SVID No delete right Nothing to recover Custody admin rogue Out-of-band RADOS ops Two-person root Custody B · tape Restore from other copy Data centre loss Site down Start control in B Custody B · OpenBao quorum Measured cross-DC RTO Restore too slow ETA > RTO in flight Escalate · pause drills Drill-measured throughput Partial, priority first Key loss Unwrap fails in drill Freeze key rotation Offline key escrow Quarterly key drill Failure Modes — Detection, Containment, What Keeps the Copies Safe Every row fails towards keeping a copy. None fails towards deleting one. v 1.0 · owner SRE · date 2026-09

Decisions

  • Every row's protection is structural: a right the identity does not have, a lock enforced by storage, a copy in a separately administered place, or a monitor outside the platform. None of them depends on someone being careful.
  • Wherever possible, detection is a rehearsal, not an alert. Key loss, version drift and a too-slow restore all appear first as a failed drill, when finding them costs nothing.

Gamedays

  • Every quarter, one whole scenario is run and timed by people following the written procedure: data-centre loss, custody credential loss, catalogue loss, or deletion of a production namespace. The measured times are published next to the targets, and the authors of the runbook are not allowed to run it.

Residual risks

  • A regional disaster that takes both data centres leaves only tape, with a multi-day RTO.
  • An application-level corruption that no invariant checks for passes depth 3 and is found only by the business. Clean-point search can recover from it, but only once someone can describe it.