On-Call Load
also called Page Volume, Interrupt Load
The count of actionable pages per shift and interrupted nights per person per quarter, which determines whether a rotation is sustainable - and which no change to the schedule can improve.
A team of six redesigns its rotation three times in a year — shift lengths, handover, the escalation policy. People keep leaving.
The schedule was never the problem. On-call viability is decided by page volume, and a rota allocates a load rather than reducing it. Six people can sustain two actionable pages a week and cannot sustain two a night, whatever the shape of the calendar.
On-call load is that volume, measured properly: actionable pages per shift, and interrupted nights per person per quarter. The first measures the alerting; the second measures the human cost, and it is the one that predicts resignations.
Why it matters
It is the only on-call metric with a feedback loop attached, and the loop is vicious. High load causes fatigue, fatigue causes attrition, attrition removes the people who knew the systems, and a rotation with less knowledge handles the same load worse. That loop, not any single bad night, is the real failure mode, and it runs over quarters so it is rarely attributed to its cause.
Load growth is also the default: every incident adds an alert and nothing removes one, so without a stated limit page volume only increases, slowly enough that each quarter feels like the last.
And measuring it makes the gap actionable. A rotation that needs two actionable pages a shift and receives eight has a measurable engineering task — fix the alerting or automate the mitigation — rather than a complaint about culture.
Implementation patterns
- Count actionable pages, not pages. Triage weekly into actionable, should-have-been-a-ticket, or false; the last two generate work items. Without the split, a fall in total pages can be entirely false-positive suppression.
- State a budget as a number. "More than two actionable pages per shift means alerting is broken, and fixing it takes priority over feature work." A budget without a stated consequence is a report.
- Track interrupted nights per person per quarter as a team metric reviewed with the manager. This is the human-cost measure, and the one that gets omitted.
- Guarantee recovery time and enforce it. Anyone paged overnight starts late or takes the day, decided by policy rather than by the individual — individual judgement under team pressure always chooses to come in.
- Route by service, with a runbook per alert naming the likely cause and the first command. The test: can a competent engineer who has never touched that service act on it?
- Automate the mitigation before staffing the response. Auto-rollback, auto-restart, automatic load shedding. A page that exists because nothing automatic happens is a page that should not exist, and this is usually the highest-return work available.
- Compensate on-call, or expect the load to be paid in attrition instead.
Industry example
Google's SRE practice made this measurable through two now-standard mechanisms: a ceiling on the share of time spent on operational work — the widely-cited 50% toil cap, with excess load pushed back to the development team — and the principle that every page should be actionable and novel. The structural idea worth copying is the feedback mechanism: giving on-call load a ceiling means exceeding it carries a defined consequence, which is what turns an observation into a control. Industry burnout research consistently finds interrupt volume and sleep disruption, rather than total hours, the strongest predictors — which is why the two metrics above are the right pair.
Failure scenarios
- Silent degradation. Load rises 15% a quarter, no quarter feels different, and within two years the rotation is unsustainable with nobody able to say when it changed.
- Measuring total pages only, so suppressing false positives looks like progress while actionable load is flat.
- Recovery time guaranteed on paper and never taken, the normal outcome without manager enforcement.
- Load concentrated on one person who is the only one able to fix a service. The rota says six; the reality is one phone.
- Alert fatigue reaching the real page. The page that mattered arrives in a channel people have learned to acknowledge without reading. This is the failure the whole practice exists to prevent.
- Volume reduced by widening thresholds. Load falls, detection falls with it, and the improvement is a reduction in coverage.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Hard page budget with an enforced consequence | Load bounded; alerting improves under pressure | Feature work stops when the budget is breached, which must be honoured |
| Business-hours-only rotation | Sustainable at small headcount; no night interrupts | Overnight failures wait, so mitigation must be automated |
| Follow-the-sun across regions | No night shifts at all | Needs headcount in multiple time zones and real handover discipline |
When not to use it
For a team of three on an internal tool, this is instrumentation nobody needs. Everyone knows how often the phone goes, and the honest move is to fix the two noisy alerts rather than build a triage ritual. Formal measurement earns its place where no single person can see the whole picture — roughly six people or two services.
More importantly, the metric is not the goal, and it is easy to improve dishonestly. Widening thresholds, muting overnight, or routing pages to a channel nobody watches all reduce the number and increase the risk. Pair every load reduction with a statement of what detection was given up, or the metric will be optimised at the expense of the thing it was protecting.
And where the honest answer is that the service should not be paged on overnight at all — a 30-minute delay costs little and an automatic mitigation exists — careful measurement just documents a rotation that should be abolished. Propose abolishing it: a tired human is a worse mitigation than a tested auto-rollback.
Interview question
Q: Your six-person team is on 24/7 for four services and two people have resigned citing on-call. The rota has been redesigned twice. What do you do?
What a strong answer covers: that the schedule is not the variable and page volume is, with actionable pages per shift and interrupted nights per quarter as the first things to measure · weekly triage into actionable / ticket / false, and why the split matters · a page budget with an enforced consequence, and the honesty that one without a consequence is a report · recovery time enforced by policy rather than individual judgement · automating mitigation ahead of staffing response · proposing business-hours-only on-call as a serious option if overnight delay is tolerable · naming the attrition loop · and refusing to cut the number by widening thresholds.
Quick check
Quiz: Which two numbers decide whether a rotation is sustainable? — Actionable pages per shift, and interrupted nights per person per quarter. Total page count is a weaker proxy for the first and says nothing about the second.
Flashcard: Why does redesigning the rota rarely fix an on-call problem? — A rota allocates the load; it does not reduce it. Viability is set by actionable page volume, so the fixes are better alerting and automated mitigation, not a different calendar.