Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
4 to work through
-
intermediate
A design platform's synchronous editing experience and its asynchronous export rendering share a compute cluster. Exports occasionally saturate it and the editor becomes unusable. What would you change?
2 min answer -
advanced
A financial platform's card authorisation path must survive dependency failures, deployments, overloaded downstreams and partial network failures. How should isolation, deadlines, breakers, bulkheads, caching, shedding and fallback combine?
2 min answer -
advanced
A security vendor pushes a content update to millions of endpoints simultaneously and a malformed file crashes them all at kernel level. What should have been in place, and why is "it was data, not code" the wrong defence?
3 min answer -
advanced Multiple choice
What is the single highest-value fault isolation mechanism, and why is it usually missing?
2 min answer
5 terms in this topic
Blast Radius
The set of users, tenants, regions or services that a single failure can affect - the quantity that distinguishes an incident from a catastrophe, and…
practiceBlast Radius Analysis
Determining, for each component, exactly what fails and who is affected when it fails completely.
patternConcurrency Limiting
Bounding the number of simultaneous in-flight operations so that overload produces fast rejection rather than resource exhaustion.
conceptFault Domain
A boundary within which a single failure is contained, defined by the infrastructure and dependencies that components inside it share.
conceptRecovery Cost Per Unit
The effort required to restore one affected instance, device or tenant - the multiplier that decides whether a wide blast radius is an incident or a …
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.