Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
6 to work through
-
advanced
A digital bank needs automatic failover for availability, but some operations cannot safely execute twice. Which components fail over automatically, which degrade, and which stop?
2 min answer -
advanced Multiple choice
A load balancer health check is changed from "the process answers" to "the process can reach its database and its cache". The dependency has a brief regional problem. What happens to the fleet, and why do load balancers deliberately fail open?
3 min answer -
advanced
A regional failover completes in 90 seconds per the runbook - the database is promoted and traffic is redirected - but the application stays broken for 25 minutes. The new region's services are healthy and idle. Where is the time going?
3 min answer -
advanced
A search platform's index-serving region becomes degraded but not fully down - elevated latency and partial errors. Should traffic fail over automatically? Analyse the risks either way.
2 min answer -
advanced
The business asks for multi-region active-active. Walk through the decision and what it actually requires.
2 min answer -
advanced
Your database primary is unreachable from the monitoring system but is still serving some clients. Do you fail over?
2 min answer
5 terms in this topic
Blast Radius Staging
Applying a change or a traffic shift in increasing increments gated on health, so that a wrong decision affects a bounded population before it affect…
patternFailover
Switching to a standby when the primary fails — where detection, fencing and the decision to automate are harder than the switch itself.
patternFailover Orchestration
The sequence of detection, decision, promotion and traffic redirection that moves service from a failed component to a healthy one.
metricFailover Propagation Delay
The interval between a failover completing on the server side and the last client actually using the new location, which is usually the larger part o…
conceptHealth Check Depth
How much of a service's dependency graph a health check consults, which determines whether the check reports an independent fault or a shared one tha…
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.