RTO & RPO
How long recovery may take and how much data may be lost.
5 to work through
-
beginner Multiple choice
A team documents an RPO of 24 hours because backups run nightly at 02:00 and take about 40 minutes. The last restore was attempted at the time the system was built. What is the honest recovery point objective?
3 min answer -
intermediate
A financial-operations platform sets one RTO and RPO for the entire business. What goes wrong, and how should they be derived instead?
2 min answer -
advanced
A collaborative workspace product must define RTO and RPO. The product team says "we can never lose a user's work". What does that requirement actually mean, and what does it cost?
2 min answer -
advanced
A stakeholder asks for zero RPO across the estate. What does that cost and what would you propose instead?
2 min answer -
advanced
In April 2022 a maintenance script at Atlassian used the wrong identifiers and deleted sites belonging to about 775 customers. Restoration took up to about two weeks for some of them, even though backups existed and were working. What property of the recovery design accounts for that gap?
3 min answer
4 terms in this topic
Recovery Point Objective
The maximum acceptable data loss measured as a duration, determining the replication and backup strategy.
metricRecovery Time Objective
The maximum acceptable duration between a failure and restored service, agreed with the business rather than chosen by engineering.
conceptRestore Granularity
The smallest unit a recovery procedure can return without disturbing everything stored alongside it, which sets the real recovery time for any incide…
metricRTO and RPO
How long recovery may take, and how much data may be lost — the two numbers from which every disaster recovery design follows.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.