Tail-Based Sampling
also called Post-Trace Sampling, Intelligent Sampling
Deciding whether to keep a trace after it completes, so that errors and slow requests are always retained while ordinary traffic is sampled cheaply.
Head-based sampling makes the decision when a request begins, before anything is known about it — so the choice is necessarily random. Tail-based sampling buffers a trace's spans until it completes, then decides based on what actually happened: it errored, it was slow, it came from a customer of interest, or it was randomly selected as a baseline sample.
The economics invert. Instead of retaining a random 1% and hoping the interesting requests are among them, you retain exactly the requests worth looking at plus a small statistical baseline.
Why it matters
Uniform random sampling fails in the specific case observability exists for. At a 1% rate, an error occurring in 0.01% of requests is almost never captured. An engineer investigating a customer's failing request will not find it, tries twice, and stops using the tool — which is how expensive tracing deployments end up with low adoption.
Rare-but-severe conditions — a failure affecting one region, one device class, or one large customer — are precisely what uniform sampling is indifferent to.
Implementation patterns
- Consistent decisions across the whole trace. Every span of a trace is kept or discarded together. A partially-sampled trace is worse than none, because it looks complete and is not.
- Record the sampling rate on the retained data, or sampled traces cannot be used quantitatively. Deriving counts from sampled data without the rate is a silent, common error.
- Compute metrics before sampling. Rates, error counts and duration histograms come from unsampled counters; sampling reduces only the volume of detailed traces. This preserves statistical accuracy while cutting cost.
- Per-route and per-tenant rate adjustment, so a low-traffic but important route is sampled heavily and a high-volume health check barely at all.
- A forced-keep mechanism — a debug header, a known-problematic tenant, an error already detected — that works even in a head-based deployment.
Industry example
Platforms handling extreme request volume at the edge cannot retain telemetry for every request, and the naive responses both fail. Retaining a uniform sample loses the interesting requests. Retaining only errors loses the baseline needed to tell whether a latency profile is unusual, because diagnosis requires contrast.
Reducing retention duration instead does not help either, because at that scale the cost is dominated by ingestion and indexing rather than by how long data is stored — and it removes the ability to investigate anything historical.
The configuration that works keeps all errors and slow requests, a small random baseline, consistent decisions per trace, and unsampled metrics alongside. The price is a collector tier that buffers every trace until completion, sized for peak throughput — which is a real infrastructure cost and the main reason head-based sampling persists.
Failure scenarios
- Inconsistent sampling across services, producing broken traces that mislead.
- Sampling rate not recorded, so all derived counts are wrong by an unknown factor.
- Metrics derived from sampled traces, which introduces error into the numbers used for alerting.
- Collector buffer overflow under load, dropping traces exactly during an incident.
- Sampling applied before the error occurs in a head-based system with no forced-keep path.
Trade-offs
Tail-based sampling requires buffering, which means memory, a collector tier and a new failure domain on the telemetry path. It also adds latency between a request completing and its trace being queryable.
Head-based sampling is simpler, cheaper and stateless, and it is a reasonable choice when paired with a generous baseline rate and a forced-keep mechanism for requests that identify themselves as interesting early. The decision turns on volume: at moderate scale, head-based with forced-keep captures most of the benefit.
Interview question
"You can only afford to store 1% of your traces. Design the sampling strategy, and tell me what breaks in your metrics if you get it wrong."