metric

On-Call Floor

also called Sustainable Rota Size, Responder Floor

The minimum number of qualified responders a rota needs before a round-the-clock commitment is staffable, which turns a published platform objective into a headcount question rather than an effort question.

on-callplatform teamstaffingsloteam topologies

A platform team of four publishes a 99.95% monthly objective on its ingress and config service. The objective is technically achievable, the components are redundant, and the commitment still fails for a reason that has nothing to do with architecture. Four people cannot hold a 24-hour rota: each is on call 13 weeks a year, and one resignation puts the remaining three on one week in three.

The on-call floor makes this visible before the commitment is published, where a platform's reliability promise stops being an engineering question and becomes a staffing one.

Why it matters

A published objective implies a human who responds out of hours, and the arithmetic is unforgiving: on a weekly primary rota of N people each person is on call 52/N weeks a year - 13 weeks at N = 4, about 9 at N = 6, about 6 at N = 8. Most organisations find a sustainable ceiling near one week in six, and below that attrition becomes the binding constraint, which removes a responder and makes the rota worse, so the failure compounds.

Product teams also read the number as a guarantee and use it in their own availability arithmetic. If the 3 a.m. response is one person who may be on a flight, their commitments rest on a false input.

Implementation patterns

  • Compute 52/N before publishing an objective and put it in the same document as the SLO.
  • Count qualified responders, not headcount. Someone who cannot safely act at 3 a.m. is not on the rota, which often puts real N two or three below team size.
  • Budget the page rate. Beyond roughly 2 out-of-hours pages in a week of primary duty, reliability work stops being optional - a better signal of team health than any survey.
  • Prefer a business-hours objective with documented fail-static behaviour when N is below the floor. "99.9% between 08:00 and 20:00 local, and here is how the client library behaves when we are unavailable" is honest and keepable.
  • Or reduce the surface: publish fewer objectives and mark the rest best-effort.

Industry example

The pattern shows up wherever a small platform group publishes a round-the-clock number: the component meets its target and the team does not survive the year, so the postmortem is about attrition rather than reliability. Published SRE practice treats on-call load as a design constraint with a page-rate budget and a minimum rota size rather than a matter of willingness, and teams running in production above the floor report a different experience from those at three or four.

Failure scenarios

  • The heroic rota, where one engineer absorbs most out-of-hours load informally, so the number looks staffed until that person leaves.
  • Silent erosion: pages are acknowledged and worked next morning, so the SLO reports met while consumers had an 8-hour outage.
  • Qualification collapse after a reorganisation, when the two people who understood the component move and nobody notices N fell to 2, with no escalation path for a primary who cannot fix it.
  • Objective inflation, where each consumer negotiates a higher number and nobody recomputes the staffing.

Trade-offs

Choose Gains Pays
Publish 24x7 below the floor Consumers get the number they asked for Burnout then attrition then a broken promise
Business hours plus fail-static terms Honest and keepable with a small team Consumers must design for degradation themselves
Pool responders across teams N rises without new headcount Less depth per responder and more escalation

When not to use it

The floor constrains published commitments, not every component. An internal tool whose unavailability costs an hour of inconvenience needs no rota, and applying this calculation to it manufactures work.

It also does not apply where the consequence waits. A batch platform whose jobs have a 24-hour deadline can be run by a small team in working hours. The question is not how important the component is but how quickly a human must act; where the answer is "by tomorrow", a rota would be theatre.

Interview question

Q: You run a platform team of five. Two product teams need 99.99% on your ingress and will fund it. What do you tell them, and what would you build?

What a strong answer covers: that 99.99% is about 4 minutes a month, so the commitment is about detection and response and therefore about a rota five people cannot staff · the arithmetic out loud, 52/N against a one-week-in-six ceiling · that the funding buys responders or isolation and probably both, since a better number on shared instances is a promise with no mechanism · the alternative of a lower published number plus fail-static terms · and measurement from the consumer's side.

Quick check

Quiz: A team of four wants to publish a 24x7 objective. Which number decides whether they can? 52/N weeks of primary duty per person - 13 weeks at N = 4 against a ceiling near one week in six, so it needs about six qualified responders or a business-hours objective instead.

Flashcard: Why is a published platform SLO a headcount decision? — It implies out-of-hours human response, and on a rota of N each person is on call 52/N weeks a year; below roughly six the binding constraint is attrition rather than architecture.