metric

Availability Denominator

also called SLO Denominator, Error Budget Denominator, Unit of Availability

The unit a nines target is computed over - calendar minutes, service-window minutes or successful requests - which changes the same percentage by an order of magnitude and decides whether the target can be measured at all.

availabilitysloerror budgetmeasurementinternal tools

A brief says 99.99%. Three people in the room agree to it, and they have agreed to three different things. Over a 30-day calendar month, 0.01% is 4.3 minutes. Over an 11-hour weekday service window it is 1.5 minutes. As a ratio of successful requests on a service doing 264,000 requests a month, it is 26 failed requests.

The percentage carries no information until the denominator is named. The argument that follows, during the first incident review, is not about the number: it is about whether a 90-second restart at 02:00 counts, and whether one user's ten retries over bad wifi are ten outages.

Why it matters

A request-ratio denominator measures user harm and a clock denominator measures service presence, and the two diverge most for the systems people set aggressive targets on: low-traffic internal tools, batch endpoints, and anything with a pronounced daily cycle.

It also decides whether the target is measurable. At 18 requests a minute, a budget of 26 failures a month sits inside the noise of client behaviour: one user retrying a form moves it, and a 20-minute outage at 06:00 does not, because nobody was asking. A metric dominated by who happened to be logged in is not a control.

Implementation patterns

  • Clock-based with a synthetic probe. Run a probe every 60 seconds against a real code path and define unavailability as consecutive probe failures. This gives a denominator independent of traffic, and it is the default for anything under a few hundred requests a minute.
  • Request-ratio with a minimum sample size. Only credible above the traffic level where a window holds thousands of requests. Below that, aggregate to a longer window or switch denominators.
  • Service-window targets. Measure 08:00 to 19:00 on weekdays and state that explicitly. This is honest for a tool nobody uses at night and it hands you the maintenance window for free.
  • Exclusions written in advance: announced maintenance, client-side network failures, requests rejected for quota. Agreed before the first incident these are policy; agreed after, they read as an excuse.

Industry example

Public cloud service level agreements are the best-documented use of the choice. Compute SLAs are generally written as a monthly uptime percentage over clock minutes per region, while API and storage SLAs are written as an error rate over billable requests. The split is mechanical: a virtual machine either exists or does not, so clock time describes it, whereas an object store serves a vast number of independent requests, so a ratio does. A probe-based denominator measures what the service does in production whether or not a customer is asking, which a ratio cannot. Read a vendor SLA for its denominator before its percentage, then read your own documents the same way.

Failure scenarios

  • The unmeasurable target. 99.99% request-ratio on a low-traffic service: the dashboard shows 100% most months and 99.2% in the month a single scripted client misbehaved.
  • The invisible outage. A business-hours denominator on a service that feeds an overnight batch, so a four-hour failure at 02:00 shows as a perfect month while a downstream deadline is missed.
  • Maintenance eats the budget. One 90-second restart a week is about 6 minutes a month, so the deployment practice has already decided the achievable number before the design starts.

Trade-offs

Choose Gains Pays
Clock minutes with a probe Traffic-independent, simple to argue, cheap Counts failures nobody experienced; a probe can miss partial failures the real traffic hits
Request ratio Tracks user harm; naturally weights busy periods Needs volume to be stable; noisy at low traffic; invites argument about which requests count
Service-window minutes Honest for office-hours systems; frees maintenance time Hides out-of-hours risk

When not to use it

Do not run this debate for a system where neither number is acted on. With nobody on call and no credit, penalty or decision attached to a breach, the honest artefact is a stated recovery time objective and a tested restore, not an availability percentage. The denominator question earns its cost the moment someone is paged, paid or penalised against the number.

Interview question

Q: A product owner asks for 99.99% availability on an internal tool used by 600 staff during office hours. Walk me through what you would agree to, and how you would measure it.

What a strong answer covers: converting the percentage under two denominators and showing they differ by an order of magnitude; identifying request-ratio measurement as useless at this traffic level; proposing clock minutes inside the service window with a 60-second probe; checking whether anything out of hours depends on the service; and noting that the weekly restart already exceeds the proposed budget, so either the target or the deployment practice has to change.

Quick check

Quiz: 99.99% over a 30-day month, over an 11-hour weekday window, and as a request ratio at 264,000 requests a month - give the three budgets. Answer: about 4.3 minutes, about 1.5 minutes, and 26 failed requests.

Flashcard: Why is a request-ratio target unusable on a low-traffic internal service? At roughly 18 requests a minute one user's retries move the monthly number while an out-of-hours outage does not, so it measures who was logged in rather than whether the service was up.