On-Call
Sustainable rotations, actionable pages and handover discipline.
5 to work through
-
intermediate
A 24-person platform group in one country pages its engineers at night. Leadership offers to fund a second site in another time zone so nobody is woken. What does follow-the-sun buy, what does it cost, and what would you check before agreeing?
3 min answer -
intermediate
A platform's on-call rotation is causing burnout - engineers are paged frequently at night and turnover is rising. What should be measured and changed?
2 min answer -
intermediate
A team is paged eleven times a week and morale is poor. What do you do first?
3 min answer -
intermediate
An on-call rotation is producing burnout and slow responses. Alert volume is high and most pages are not actionable. What changes, and in what order?
2 min answer -
intermediate
You join a team of 12 engineers taking around 40 pages a week, of which perhaps 5 required action. The team is exhausted and two people have resigned. The director asks for a plan. What do you do, and in what order?
3 min answer
3 terms in this topic
On-Call
The rotation that responds to production problems — a system whose health is measured by whether the people in it can sustain being in it.
metricPage Budget
An explicit ceiling on how many pages a shift may generate, treated as a limit the team manages against rather than as an outcome it observes.
practiceSustainable On-Call
An on-call arrangement whose alert volume, rotation size and compensation allow it to continue indefinitely without degrading the people in it.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.