A live-streaming platform of Twitch's shape needs usable latency quantiles for 300 API endpoints whose responses span 2 ms to 90 seconds. The current histograms use 12 fixed buckets topping out at 10 seconds and every p99 above that reads as the overflow bucket. Which change fits the problem?
Show the full answer Hide the answer
The deciding property
The precision this workload needs is relative, not absolute. A 2 ms endpoint needs millisecond resolution; a 90-second endpoint does not care about 50 ms either way. Fixed buckets impose one absolute grid on both, and no single grid is simultaneously fine at the bottom and affordable at the top.
Why exponential histograms
Boundaries are successive powers of base = 2^(2^-scale), so the relative error is (base-1)/(base+1) at every magnitude — about 4.3% at scale 3. OpenTelemetry SDKs start at a maximum scale (default 20) and halve the scale, merging adjacent bucket pairs, whenever the observed range would exceed the maximum bucket count (default 160). That 160 was chosen to cover 1 ms to 100 s at under 5% relative error, which is this problem stated as a design goal. The representation maps one-to-one onto Prometheus native histograms, so the storage side exists.
The cardinality effect is the reason it matters here. A classic histogram costs one series per bucket boundary plus sum and count: 300 endpoints by 3 status classes at 12 buckets is about 12600 series, and the same coverage on a 160-bucket fixed grid would be roughly 146000. The exponential form carries its buckets inside one sample, so the same 900 label sets cost 900 series. Cost moves from series count into sample size and into needing a backend, query layer and dashboard stack that support the type.
Why the other options fail
Forty more fixed buckets applies a linear remedy to an exponential range. Holding 5% relative error from 2 ms to 90 s — a factor of 45000 — needs roughly ln(45000)/ln(1.05), about 220 buckets. Fifty-two buckets is still unusable at both ends and multiplies series by four.
Instance-computed quantiles as gauges cannot be aggregated. Each of several hundred instances reports its own p99 and there is no arithmetic that combines those into a fleet p99, because the distributions were discarded at the source. It also breaks every sum by and rate a dashboard wants.
Raising the top boundary to 120 seconds fixes overflow and not resolution. Everything between 10 s and 120 s lands in one bucket, so a slow endpoint's p99 is reported as "somewhere in a 110-second window", and the fast endpoints are exactly as coarse as before.
Quantiles from the trace store are exact and expensive: a dashboard refresh becomes a scan over raw events, and if traces are sampled the quantile inherits the sampler's bias. It is the right tool for an investigation and the wrong one for a panel that refreshes every 30 seconds.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| The backend cannot store native or exponential histograms | Two fixed grids split by endpoint class | Correctness beats elegance and a fast grid plus a slow grid covers the range |
| A contractual SLA needs exact quantiles | Compute from unsampled event records | Both histogram forms are approximations by construction |
| Endpoints all sit within one order of magnitude | Fixed buckets around the objective | Twelve well-chosen boundaries are simpler and universally supported |
When not to change anything
If the only question the metric answers is "what fraction was under 300 ms", a fixed bucket boundary placed exactly at 300 ms answers it without approximation, while an exponential histogram's answer at that point is interpolated. SLO-threshold reporting is the case where the crude scheme is the more accurate one, and swapping it out to gain range you do not use is a loss.