intermediate 4 min answer Multiple choice

A 400-service mesh runs Envoy sidecars, the data plane Lyft built and open-sourced in 2016. Tracing is on at 100% in staging and 0.1% uniform in production, and every incident review ends with "we did not have a trace for the failing request". The tracing backend can afford about 2% of production spans. Which sampling design do you adopt?

lyftenvoytracingsamplingservice meshtail-based
Pick one
Show the full answer Hide the answer

The deciding property

One fact in the stem settles it: the requests they need are the ones they cannot identify at the start. A failing request looks exactly like a succeeding request when the first span opens. Head-based sampling, by definition, decides before the outcome exists.

Everything else is detail. The question is whether the sampling decision can be made with knowledge of the outcome, and only tail-based sampling can be.

Why tail-based wins here

The collector holds the spans of an in-flight trace in memory, and when the trace completes it applies a policy to the finished object: keep if any span errored, keep if total duration exceeded the SLO, keep if it touched a service on the watch list, otherwise keep one in N. The 2% budget is then spent almost entirely on traces that are worth looking at, because the uninteresting 98% is the part being discarded.

At 2% effective retention the team keeps every error and every slow request, and those are typically well under 1% of traffic. The budget does not constrain the interesting traffic at all — only how much routine traffic is kept as a baseline, and a baseline needs far less than 2%.

What it costs

Be specific about the bill, because this option is not free:

  • Buffer memory at the collector. Every in-flight trace is held until it completes. Rough sizing: traces per second × mean trace duration × spans per trace × span size. At 20,000 traces/s, 400 ms mean duration, 25 spans, 400 bytes a span, that is 20,000 × 0.4 × 25 × 400 ≈ 80 MB of steady-state buffer, which is small. At a 30-second p99 duration for the long tail, the tail dominates the buffer and it is much larger. Size the buffer from the p99 trace duration, not the mean.
  • Span-aware routing. Every span of a trace must reach the same collector instance, so the tier needs consistent hashing on trace ID. Forgetting it is the usual reason a first rollout produces broken half-traces.
  • A hard timeout. A trace whose root never closes must be evicted or the collector leaks. Set it above p99 request duration and alert on the eviction rate.

Why the other options fail

  • Raise uniform sampling to 2%. This is the option most teams pick, and it is proportional: it keeps 2% of successes and 2% of failures. If the failing condition affects 50 requests an hour, the team retains one of them an hour, which is still "we did not have a trace". It buys a 20x increase in cost for a 20x increase in the chance of coincidence. The uniform rate is right only when the thing being studied is common.
  • Per-route head-based rates with errors at 100%. The closest wrong answer. It works for failures the service knows about and fails for the one in the stem: a request that returns HTTP 200 and is wrong, or is slow in a downstream the entry service cannot see. The decision is made at the root span, before any downstream has reported anything. A real improvement, and not the answer when failures are discovered late.
  • 100% at the edge with independent downstream decisions. This produces partial traces, which are the worst outcome: a trace that shows six of nine hops looks complete and quietly omits the hop that was slow. Sampling decisions must propagate with the context so that a trace is kept whole or discarded whole. Envoy propagates the decision in the trace context precisely so this does not happen; overriding it per service defeats the mechanism.

When this is the wrong answer, and what would flip the decision

Tail-based sampling is operational machinery: a stateful tier, a hash-routing policy, an eviction timeout and a memory budget to watch. For a mesh small enough that an engineer can hold the call graph in their head, that machinery costs more than it returns, and per-route head-based sampling with errors at 100% gets most of the value for a configuration change. The threshold is roughly where a single trace crosses more services than one team owns, because that is where "ask the team that owns the slow hop" stops being a viable substitute for a trace.

If this changes Choose Because
The mesh is 15 services, not 400 Per-route head-based The buffer tier is operational work that a small mesh does not repay
Traces must leave the cluster under a data-residency rule Head-based Buffering raw spans centrally concentrates data a regulator may object to
p99 request duration is minutes (batch, long polling) Head-based plus error sampling Buffer cost scales with duration and becomes the dominant line item
The team already runs wide events per request Tail-based, lower budget Events carry the per-request facts; traces are needed only for the causal shape

What a strong answer adds

Record the effective sample rate on every retained trace. Without it, every count derived from traces is silently wrong, and a team that has spent six months reasoning about "how often this happens" from sampled data will be out by whatever the policy's ratio happened to be that week.