advanced 3 min answer

A commerce API serves 2000 requests per second and keeps 1% of traces. The team wants to watch a 0.5% checkout error rate from the trace store and page when it moves by a tenth of itself. Roughly how much traffic must the sample cover before that reading is trustworthy, and what does the answer say about the alert?

samplingstatisticstracingalertingopentelemetry
Show the full answer Hide the answer

The assumptions, stated

Error rate p = 0.005. Sample rate 1%. Offered load 2000 rps, so the sample carries 20 requests per second. The team wants a 95% reading good to ±10% relative — that is, 0.5% ± 0.05 points, which is what "a tenth of itself" means.

The arithmetic

The precision of a proportion depends on the count of the rare thing, not on the size of the sample. For an observed count k of errors, the relative standard error is about 1/√k, so a 95% interval of ±10% relative needs 1.96/√k = 0.1, or k ≈ 400 error events inside the sample.

At p = 0.005, collecting 400 sampled errors means 400 / 0.005 = 80000 sampled requests. At 20 sampled requests per second that is 4000 seconds, or a little over an hour per reading.

The same formula run backwards is the useful version: k ≈ 100 gives ±20% relative, k ≈ 25 gives ±40%. Below roughly 25 events of the class you care about, the number is noise wearing a decimal point.

The number, and which assumption dominates

An hour per trustworthy reading, with a range of maybe 40 minutes to 2 hours depending on how much the error rate itself moves. The dominant term is p, not the sample rate: halving the error rate you are watching doubles the traffic required, while halving the sampling rate also doubles it. Traffic volume enters only through how fast that traffic arrives.

What the number rules out

A rate computed from sampled traces cannot back a page. An alert wanting a five-minute detection window would see about 6000 sampled requests and 30 errors — a reading with a ±36% relative interval, which flaps between 0.32% and 0.68% while nothing changes.

The split that works: rates come from unsampled counters, examples come from sampled traces. A counter increment costs a few bytes of a pre-aggregated series and is exact at any rate, so the SLI and the alert are computed there. The trace store then answers "show me one of those failures", which needs a handful of retained traces, not statistical power. Attach the trace id to the metric point as an exemplar so the jump from the alert to the example is one click.

If the trace store genuinely must carry the count, retain the rare class whole: tail-based sampling at 100% of errors and 0.1% of successes keeps every one of the 10 errors per second and costs about a tenth of the previous volume.

The detail that makes re-weighting possible at all

Sampled counts can only be scaled back up by 1/probability if the probability is recorded and identical across the whole trace. OpenTelemetry's consistent probability sampling puts a 56-bit rejection threshold in the th field of tracestate alongside a randomness value, so every service in a request evaluates the same inputs and downstream systems can recover the effective rate. Without a propagated threshold, services each pick their own rate, traces come out fragmented, and no re-weighting is defensible — the bias is unknown rather than merely large.

When this analysis is the wrong one

If the class you are watching is common, sampling is fine. Watching overall p99 latency at 1% of 2000 rps gives 72000 samples an hour and a very stable quantile, because the statistic depends on the bulk of the distribution rather than on a rare tail. The floor bites only for rare events, and the test is always the same: count the events, not the requests.