A team of four engineers owns one service with a 24/7 paging rota. Leadership wants them to take on two more services of similar size. Which fact about the rota decides whether that is possible?
Show the full answer Hide the answer
The mechanism
A 24/7 rota of four people means each engineer is on call one week in four, carrying 168 hours in which they must be reachable and able to act. The cost of that week is not the hours on the pager; it is what a page does to the hours around it.
An out-of-hours page costs far more than the time spent fixing it. Thirty minutes of response typically consumes the rest of the night and a large part of the following day, because interrupted sleep does not refund. Industry practice converges on a sustainable load of roughly one to two out-of-hours pages per on-call shift. Beyond that, people stop responding carefully, start acknowledging and going back to sleep, and eventually leave.
Now do the arithmetic. If the existing service pages 1.5 times a week out of hours, three similar services page about 4.5 times a week. The same four people, the same one-in-four rotation, and three times the interruption rate per shift — which is above the sustainable band, so the constraint binds before any code is written.
Attrition closes the loop. Lose one engineer and four becomes three: one week in three. A rota that is already painful degrades fastest exactly when it is most loaded, which is how teams go from stretched to unstaffed in a quarter.
Why the other options fail
- Lines of code and repository count. Size correlates weakly with operational load. A large, stable batch service can page once a quarter while a small, user-facing service with a flaky dependency pages nightly. Sizing ownership by code volume is the most common version of this mistake, and it is why "they are only small services" keeps being said about teams that are drowning.
- A shared programming language. It lowers the cost of changing the services and barely touches the cost of running them. Language uniformity is worth having and it is not the binding constraint here.
- Runbooks exist. Runbooks cut the time per page and the dependence on one expert, which is real and worth doing. They do not cut the number of times someone is woken. A good runbook turns a 2-hour page into a 20-minute page and still ruins the night.
The consequence people miss
The right response to "take two more services" is usually not "no". It is a counter-offer with the arithmetic attached: reduce the page rate first, or change the rota shape. Concretely — drive the existing service's out-of-hours pages below one per week by fixing the top two alert sources; or merge with another team to reach a one-in-six or one-in-eight rotation; or pair with a team in another timezone so nobody is woken at all. Each of those is an architectural or organisational change with a measurable effect on the same number.
When this is not the binding constraint
When there is no 24/7 obligation. An internal tool whose users are all in one timezone, or a service whose failure can wait until morning because the business process it serves does not run overnight, needs no out-of-hours rota at all. Deciding that explicitly is cheaper than every other fix, and teams rarely ask the question. The service's business criticality, not its existence, is what earns a pager.