Backup and Restore Service · View 21 of 26 · 6 · Operations
Decisions
- The verifier's liveness is watched by Healthchecks, run by the observability team outside this platform. Each tier and each Tier 1 datastore has a check with a grace period of twice its interval. Silence pages.
- A custody auditor compares every bucket's Object Lock configuration against the declared policy every hour and samples checksums. A lock quietly changed from compliance to governance is a finding, even if no data is touched.
- Every failed check is stored as a durable finding in the catalogue, not just a transient alert, so a checksum mismatch cannot be acknowledged and forgotten.
Numbers
- Protection status is at most 5 minutes stale. A verification result is visible within 15 minutes of the rehearsal finishing.
- Showback: storage priced at the amortised cost per TB-month of each custody class, and rehearsal priced by NVMe node-hours. Verification cost is reported on its own line.
Omitted
- Standard service metrics (latency, errors, saturation) for the platform's own services follow the house SLO pattern and are not repeated here.