A stakeholder writes "99.99% availability" into the brief for an internal expense-approval tool used by about 600 staff on weekdays between 08:00 and 19:00. Work out what that allows under three different denominators, and say which number you would put in the agreement.
Show the full answer Hide the answer
The assumptions, stated
600 staff, roughly 20 interactions each per working day, so about 12,000 requests a day and around 264,000 a month. The service window is 11 hours a day across about 22 weekdays, so 242 hours or 14,520 minutes a month. A calendar month is 43,200 minutes. The platform team does one rolling restart a week, each taking about 90 seconds.
The arithmetic
The same percentage produces three different budgets, and nobody in the room has said which one they mean.
- Calendar minutes. 0.01% of 43,200 minutes is 4.3 minutes a month. 99.9% would give 43 minutes; 99.5% would give 3.6 hours.
- Service-window minutes. 0.01% of 14,520 minutes is 1.5 minutes a month, because the nights and weekends you were not going to be measured on have been removed from the denominator. The stricter-looking number is the same promise.
- Request ratio. 0.01% of 264,000 requests is 26 failed requests a month. At roughly 18 requests a minute during the window, a single 90-second restart that drops connections costs about 27 requests, so the weekly maintenance alone exhausts the month's budget.
Which denominator to use
At 18 requests a minute a request-ratio SLO is statistically noisy: one user on bad hotel wifi retrying a form ten times moves the monthly number, and a genuine 20-minute outage at 06:00 moves it not at all. Low traffic breaks ratio measurement in both directions.
For a service of this size, measure minutes of unavailability inside the service window, from a synthetic probe running every 60 seconds, and write the target as 99.5% of the window, which is 72 minutes a month. The probe gives you a denominator that does not depend on whether anyone happened to be logged in.
What the number rules in and out
4.3 minutes a month rules out: any maintenance that drops traffic, single-instance deployments, schema changes that need a lock, and a dependency on anything with a lower target than yours. It also rules out learning about failures from users, because the whole budget is shorter than the time it takes a human to notice, decide and act. Buying it means redundancy, automated failover and tested rollback for a tool that 600 people use between breakfast and dinner.
When this is the wrong answer
If the tool sits on the critical path of something that does not keep office hours - a payroll cutoff, a regulatory filing window, an overnight batch that expects the API - the service window is the wrong denominator and a calendar target is correct. Ask one question before choosing: what is the worst hour for this to be down, and is that hour inside the window? If the answer is outside it, the window is a measurement convenience that hides the risk you actually care about.