Low-Traffic Alerting
also called Small-Denominator Alerting, Sparse-Signal Alerting
The set of techniques that keep ratio-based alerts meaningful when request volume collapses overnight or on a low-traffic service - where two failures out of forty is five percent and pages someone for noise.
A checkout service does 4,000 requests a minute at midday and 40 at 03:00. The error-rate alert is set at 5%, which is a careful, well-argued threshold derived from the SLO. At 03:00, two failed requests is 5%, and two failed requests is what a flaky retry, a single bad client or an ordinary timeout produces on any given night.
So the alert fires. It fires most nights. Within a month the responder has learned to acknowledge it without reading it, and the mechanism that was supposed to protect the service has instead trained the only person who could save it to ignore its pages.
This is not a threshold that was set badly. It is a ratio evaluated over a denominator that has collapsed, and no choice of percentage fixes it: raise the threshold to 20% and you need eight failures at night but have lost all sensitivity during the day.
Why it matters
The variance of an observed rate scales with the inverse square root of the sample count. A 1% true error rate observed over 4,000 requests sits in a narrow band; the same 1% rate observed over 40 requests produces zero, one or two errors routinely, which reads as 0%, 2.5% or 5%. The signal has not changed. The measurement has become too coarse to say anything, and the alert is reporting the coarseness.
Two failure directions follow, and teams usually only defend against the first:
- False pages at low volume, which is the dominant source of 03:00 noise and therefore the dominant cause of alert fatigue.
- Missed outages at low volume. If traffic falls to zero because the load balancer is broken, there are no requests, so there are no errors, so the error rate is 0% and nothing fires. A total outage can look identical to a quiet night. This is the more dangerous of the two and it is almost always discovered the hard way.
Implementation patterns
- Require a minimum event count. The ratio alert cannot fire unless the denominator exceeds a floor — say 100 requests in the window. Below that it is suppressed, not evaluated and ignored. This is the single highest-value change and it takes one line in most alerting languages.
- Pair every ratio with an absolute floor alert. "Fewer than N successful checkouts in 10 minutes" works precisely in the regime where the ratio cannot, and it is what catches the zero-traffic outage. The two alerts are complements, not alternatives, and having only the ratio is the common mistake.
- Lengthen the window instead of raising the threshold. Thirty minutes at 40 requests a minute is 1,200 events, which is enough. The cost is detection latency, usually an acceptable trade overnight.
- Use multi-window burn rates. Evaluating a fast window and a slow window together, as SLO burn-rate alerting does, naturally handles volume variation because it measures budget consumption rather than an instantaneous rate.
- Alert on the absence of data explicitly. Most systems treat "no data" as "not firing". Configure a separate no-data condition with its own routing, or the quietest possible failure is also the loudest possible outage.
- Scale the expectation. Compare against the same hour last week, so a drop to 40 requests is only interesting if last Tuesday had 400.
Industry example
This is why SLO-based alerting, popularised by Google's SRE practice and now standard across the industry, is expressed as error-budget burn rate rather than instantaneous error rate. A burn rate asks "at the current pace, how long until the month's budget is gone?" — a question about accumulated events over a defined window, and therefore far less sensitive to a momentary collapse in volume than a point-in-time ratio. The multi-window, multi-burn-rate formulation exists precisely to stay sensitive to a fast catastrophic burn while ignoring the small-sample noise a single short window produces.
Failure scenarios
- Chronic nightly false pages, leading to acknowledged-without-reading behaviour and a rotation nobody will join.
- The zero-traffic blind spot. Upstream routing breaks, traffic goes to zero, error rate is 0%, every dashboard is green, and the first signal is a customer.
- Threshold inflation. The percentage is raised each time it misfires, until the alert only fires for outages that were obvious anyway.
- Suppression that is never lifted. A "mute overnight" stopgap becomes permanent, leaving the service unmonitored for a third of every day.
- Per-instance evaluation at low volume. Splitting already-sparse traffic across 20 instances makes every instance's denominator tiny, so the noise multiplies by the fleet size.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Minimum event count | Removes nearly all small-sample false pages | A genuine low-volume failure is invisible until the floor is reached |
| Longer evaluation window | Statistically sound without losing sensitivity | Slower detection, which costs money on a high-value flow |
| Absolute floor alert | Catches the zero-traffic case the ratio cannot | Needs a volume expectation per service, which drifts and must be maintained |
When not to use it
On a consistently high-volume service, this machinery is unnecessary complexity. If the denominator never falls below tens of thousands per window, a plain ratio is statistically sound at all hours, and adding minimum counts and floor alerts gives three things to maintain and one more way to misconfigure a page. Add them when the daily peak-to-trough ratio is large — roughly an order of magnitude or more — or when the service is genuinely low-volume at all times.
And if the underlying question is simply "did anything break", a synthetic probe beats any amount of statistics on sparse real traffic. A request issued every 30 seconds from outside gives a constant, known denominator and detects the zero-traffic outage directly. For a low-volume but business-critical endpoint, one synthetic check replaces most of the list above.
Interview question
Q: Your error-rate page fires most nights and has never once corresponded to a real problem, but you are not willing to delete it. What do you do?
What a strong answer covers: identifying the denominator rather than the threshold as the cause, with the variance argument stated plainly · a minimum event count as the immediate fix · recognising unprompted that the same alert is blind to a zero-traffic outage, and adding an absolute floor to cover it · the window-lengthening trade between statistical soundness and detection latency · burn-rate alerting as the structural answer · and the observation that a synthetic probe may make the whole problem go away for a low-volume endpoint.
Quick check
Quiz: Why can a 5% error-rate alert be correct at midday and useless at 03:00 on the same service? — The threshold is a ratio; at 40 requests, two ordinary failures is 5%, so the alert reports sampling noise rather than a change in the service.
Flashcard: What does a ratio alert fail to detect when traffic falls to zero? — Everything. No requests means no errors means 0%, so a total outage is indistinguishable from a quiet night. Pair every ratio alert with an absolute floor or a no-data condition.