Postmortems
Finding the systemic conditions, and completing the actions afterwards.
4 to work through
-
intermediate
A platform publishes detailed public postmortems for significant incidents. What does this practice cost, and what does it buy that internal postmortems do not?
2 min answer -
intermediate
An engineer ran a command that deleted production data. What does the postmortem investigate?
2 min answer -
advanced
An engineer runs a routine capacity-removal command with a typo. It removes far more than intended and a core service is down for hours. What does the postmortem conclude?
2 min answer -
advanced
An organisation writes thorough postmortems and keeps experiencing the same class of failure. What is missing?
2 min answer
3 terms in this topic
Blameless Postmortem
An incident analysis structured so that participants can disclose what actually happened without personal risk - which is a technical requirement for…
practiceContributing Factors Analysis
Identifying the multiple conditions that combined to produce an incident, in place of searching for a single root cause.
practicePostmortems
Structured learning from an incident, conducted so that the truth is obtainable — which requires that telling it is safe.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.