advanced
4 min answer
"You are on-call for a checkout service. Design what pages you at 03:00." That is the whole prompt. How do you handle it?
Show the full answer Hide the answer
What the interviewer is testing
Whether you reach for signals or for consequences. A weak answer lists metrics; a strong answer starts from "what is the user unable to do, and does a human need to act in the next fifteen minutes?" The prompt is deliberately underspecified because the scoping is the assessment.
A secondary test: whether you treat the alert as a technical artefact or as a contract with a person whose sleep it is spending.
The clarifying questions that change the answer
Ask these and the shape of the answer changes substantially:
- What is the SLO, and is there one? Without a target, "slow" is undefined and every threshold is arbitrary. If the answer is "we don't have one", say that defining it is step one, because the alert threshold is derived from it rather than guessed.
- What is the traffic profile overnight? A checkout service doing 4,000 requests/minute at noon and 40 at 03:00 cannot use the same error-rate threshold: at 40 requests, two failures is 5% and is probably noise. Low-traffic alerting is a genuinely different problem, and knowing that is a senior signal.
- What can be fixed at 03:00? If the only remediation is "wait for the payment provider", a page is a notification with extra cruelty. Route it to a dashboard and page on the things a human can act on.
- Is there a business cost per minute? Checkout usually has one, and it is what justifies a page at all, so it should be the thing used to set the urgency tiers.
- Who else is paged by the same event? Four teams woken for one incident is a design flaw.
A strong answer's arc
Page on symptoms, at the SLO boundary, with a burn rate.
- The primary page is the SLO burn rate, not the error rate. If the monthly budget is 0.1% of requests, alert when the recent burn rate would exhaust the budget early: a fast burn (for example, 14x over an hour) pages immediately; a slow burn (2x over six hours) opens a ticket. This single mechanism replaces most threshold alerts and it self-adjusts for traffic volume, because it is a ratio against a budget rather than an absolute count.
- Two or three symptom alerts for things the SLO cannot see. Checkout-completion rate against the same hour last week catches the case where every request returns 200 and no order completes — the failure that no technical metric detects. Payment-authorisation success rate by provider. Order-to-settlement backlog.
- Low-traffic handling, explicitly. Require a minimum event count before the ratio alert can fire, and pair it with an absolute floor alert ("zero successful checkouts in 10 minutes") that works precisely when the ratio does not. A ratio over a tiny denominator is the most common source of 03:00 false pages.
- Everything else is a ticket. CPU, memory, disk, queue depth, individual pod restarts. They are inputs to diagnosis, and none of them is evidence that a user is affected.
- Every page carries a runbook link, the likely cause, and the first command to run. If the responder has to think about where to start, the alert is half-built.
Common weak answers
- Alerting on causes rather than symptoms. "Page if CPU > 80%." The user does not experience CPU. High CPU with a healthy service is a false page; a broken service at 20% CPU is a miss. Both happen weekly.
- A threshold per dependency. Nine services, each with latency and error alerts, means one real incident fires fourteen pages and the responder spends the first ten minutes working out which one is the cause.
- "Page on anything anomalous." Anomaly detection without a consequence model pages on Black Friday, on the deploy, and on the clocks changing.
- No mention of what stops paging. An alert with no auto-resolve and no silence mechanism trains people to ignore the channel, and alert fatigue is the failure mode that makes every other alert worthless.
What a strong answer adds
Three things that separate a staff-level answer:
- A budget for pages, stated as a number. "More than two pages per on-call shift means the alerting is broken, and fixing it is prioritised as production work." Without a number, alert volume only ever grows.
- The review loop. Every page is triaged afterwards into actionable, should-have-been-a-ticket, or false. The last two categories generate work items. Alerting is a system that decays without maintenance, and saying so is the difference between having designed alerts and having designed an alerting practice.
- The human cost as an architecture constraint. A rotation that cannot be sustained produces attrition, and attrition produces the loss of the knowledge the rotation depends on. That is an architectural consequence of an alerting decision, and naming it is what makes the answer senior rather than thorough.