Incident Management
Command roles, severity levels and mitigation before diagnosis.
8 to work through
-
intermediate
A payment platform is mid-incident: transactions are partially failing, the cause is unclear, and several teams are investigating. What structure makes the response effective?
2 min answer -
intermediate Multiple choice
An incident is underway and everyone is debugging. What role is missing?
1 min answer -
intermediate
An observability platform's incident process treats all incidents with the same severity and process. What problems does this create, and how should severity be defined?
2 min answer -
intermediate
You are incident commander. Error rate is 8% and rising, cause unknown. What do you do in the first ten minutes?
2 min answer -
advanced
A 43-second network blip triggers automated database failover across regions. Service is degraded for over 24 hours. Analyse.
2 min answer -
advanced
A configuration change disconnects a company's internal network. Engineers cannot access the tooling needed to revert it. Analyse.
2 min answer -
advanced
A post-mortem finds the telemetry system depended on the same infrastructure that failed, leaving engineers blind. What must be isolated, and what is the minimum set of break-glass signals?
3 min answer -
advanced
Incidents at your company are chaotic: unclear ownership, no communication, and postmortems that produce nothing. Design the improvement.
2 min answer
3 terms in this topic
Incident Severity Levels
A small, agreed scale of incident severity that determines response, escalation and communication without requiring debate during the event.
practiceRunbook Quality
The properties that make an operational procedure usable by a tired responder under pressure, as opposed to a document that merely exists.
metricTime to Detect
The interval between a problem beginning and someone knowing about it, which is often the largest and most reducible component of total incident duration.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.