DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
5 to work through
-
intermediate
A core internal service has run continuously for three years. What risk does that create?
2 min answer -
intermediate
An enterprise runs an annual disaster recovery test that always succeeds, yet the team has low confidence in real recovery. What is likely wrong with the test?
2 min answer -
advanced
An auditor requires evidence that a four-hour RTO is achievable for a system that has never failed over. Production cannot be risked and the business will not accept an unplanned outage. Sequence the first real disaster-recovery test.
3 min answer -
advanced
In January 2017 GitLab lost roughly six hours of database data after an engineer deleted a directory on the wrong host during replication troubleshooting, and then found that several backup and replication mechanisms had silently not been working. What signal would have revealed the broken backups beforehand, and why did nothing report them?
3 min answer -
advanced
Your DR plan is pilot light with a 30-minute RTO. What would you test, and what will the test probably reveal?
2 min answer
2 terms in this topic
Dark Failover
Bringing a standby environment to full serving readiness and exercising it with synthetic or internal traffic while production continues untouched, s…
practiceRestore Verification
Periodically performing a real restore from backup and validating the result, as the only evidence that a recovery capability exists.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.