advanced
4 min answer
"Design the on-call rotation for a team of six running four services." That is the prompt. How do you handle it?
Show the full answer Hide the answer
What the interviewer is testing
Whether you treat on-call as a scheduling problem or as a system with a load, a capacity and a failure mode. The weak answer produces a rota. The strong answer asks what the page volume is, because at six people the rotation's viability is decided by that number and by nothing in the schedule.
It is also a test of whether you will say the uncomfortable thing: that a six-person team may not be able to sustain 24/7 on-call at all, and that the honest answer might be to change the scope rather than the shifts.
The clarifying questions that change the answer
- How many pages per shift, and what fraction are actionable? This is the first question and it determines everything. Two pages a week is a sustainable rotation; two pages a night is an attrition mechanism.
- Are the services customer-facing, and what is the cost of a 30-minute delay? If overnight failures cost little, business-hours-only on-call with a next-morning queue is the correct design, and proposing it is a strong signal rather than a weak one.
- Is anyone the only person who can fix a given service? A rotation with a single point of human failure is a rotation that resolves to one person's phone.
- What is the time-zone spread? Six people in one city cannot follow the sun; six across two regions can, and that changes the design completely.
- Is on-call compensated, and is recovery time guaranteed? If not, the rotation has a hidden cost that will be paid in resignations.
A strong answer's arc
- Primary and secondary, one week each, so one person in six is primary, with secondary offset by a week rather than shadowing — the secondary should not be the person who was just primary. At six people this gives roughly a six-week cycle, which is about the floor for sustainability.
- Guaranteed recovery. Anyone paged overnight starts late or takes the day. Written down, enforced by the manager, not left to individual judgement, because individual judgement under team pressure always chooses to come in.
- A page budget as an explicit target. "More than two actionable pages per shift means alerting is broken, and fixing it takes priority over feature work." Without a number, page volume only grows.
- Weekly triage of every page into actionable, should-have-been-a-ticket, or false, with the last two categories generating work. This is the feedback loop that keeps the rotation viable, and it is the first thing dropped under delivery pressure.
- A runbook per alert, linked from the page, naming the likely cause and the first command. The test is whether a competent engineer who has not touched that service can act on it.
- Explicit escalation, with service owners reachable as a documented second tier, so the primary's job is triage and mitigation rather than deep expertise in four services.
- Routing by service, not by severity. Four services and six people means nobody is expert in all of them; the page should carry which service and which playbook.
Common weak answers
- A rota and nothing else. The schedule is the easy part and is not the thing that fails.
- No page budget. Alert volume grows without a stated limit, and the rotation degrades quarter by quarter until people leave.
- Shadowing as onboarding and nothing more. Shadowing builds familiarity and not the confidence to act alone at 03:00. That needs game days and a first shift with a named, awake buddy.
- Ignoring the arithmetic. Six people, 24/7, with high page volume does not work, and a candidate who does not say so has not answered the question.
- Treating documentation as the fix for everything. Unfalsifiable and unmeasurable.
What a strong answer adds
- The honest option of not doing 24/7. If overnight failures are tolerable for 30 minutes, a business-hours rotation with an automated mitigation path (auto-rollback, auto-restart, load shedding) is cheaper, kinder and more reliable than a tired human. Automating the mitigation is usually a better investment than staffing the response.
- On-call load as a measured quantity with an owner. Pages per shift, interrupted nights per person per quarter, and time-to-acknowledge, reviewed as a team metric. A rotation nobody measures is one that degrades invisibly.
- The attrition argument, stated plainly. A rotation that burns people out loses the knowledge the rotation depends on, which makes the next rotation worse. That feedback loop is the real failure mode, and naming it is what makes the answer senior rather than organised.