pattern

Head-Based Sampling

also called Upfront Sampling, Root-Span Sampling

Deciding whether to keep a trace at the moment its first span opens and propagating that decision to every downstream service, so cost is bounded at the source at the price of never being able to select on the outcome.

tracingsamplingtelemetry costenvoylyftpropagation

Every incident review ends the same way: "we did not have a trace for the failing request." Tracing is on, the budget is spent, and the one request anybody wanted is the one that was discarded. The reason is structural, not a misconfiguration.

Head-based sampling makes the keep-or-drop decision at the root span, before the request has done anything. The decision is written into the trace context — the traceparent header's sampled flag — and every downstream service honours it rather than deciding again. That propagation is the point: it is what makes a trace whole or absent rather than partial.

The consequence is unavoidable: a failing request is indistinguishable from a succeeding one at the moment the decision is made. Sampling at the head cannot select for outcomes, because outcomes do not exist yet.

Why it matters

It is the default in every tracing SDK, and for the majority of systems it is the correct default. It is the only sampling strategy with no infrastructure: no buffer, no stateful tier, no routing constraint, and a cost that is exactly the configured fraction of traffic.

It also has a property tail-based sampling lacks: the saving happens at the source. A dropped trace is never serialised, never sent and never paid for in network or collector capacity. Tail-based sampling must transport and buffer everything before discarding most of it, so its cost saving is only at the storage layer.

Knowing where its ceiling is matters as much as knowing how it works. The ceiling is reached the moment the question becomes "show me the request that failed".

Implementation patterns

  • Uniform probabilistic. Keep 1 in N of everything. Simple, unbiased, and the wrong tool for rare events: at 2%, a condition affecting 50 requests an hour yields one trace an hour.
  • Per-route or per-service rates. A high-volume health endpoint at 1 in 10,000, a checkout flow at 1 in 10. This is the highest-value refinement available and is a configuration change.
  • Rate-limited sampling. "At most 100 traces per second per service" bounds cost under a traffic spike, where a percentage does not. Pair it with a probabilistic floor so low-traffic services still get coverage.
  • Error-aware at the root. Keep 100% of requests the entry point already knows have failed — a rejected auth, a 4xx at the gateway. This captures the failures visible at the front door and nothing discovered later.
  • Debug override. A header or flag that forces sampling for a specific request, so support and engineers can trace a reported problem on demand. This single feature recovers much of what head-based sampling gives up, and is frequently omitted.
  • Always record the effective rate on the trace. Without it, every count derived from traces is silently wrong by the policy's ratio.

Industry example

Envoy, the data plane Lyft built and open-sourced in 2016, propagates the sampling decision in the trace context across every hop in a mesh. The design choice worth noting is that services are expected to honour the inherited decision rather than make their own. A service that re-decides produces a trace showing six of nine hops — and a partial trace is worse than no trace, because it looks complete while omitting the hop that was slow. Centralising the decision at the edge and propagating it is what keeps traces whole, and it is why the sampling rate in a mesh is a platform setting rather than a per-service one.

Failure scenarios

  • The rare failure is never captured, so the mechanism fails precisely for the question it was bought to answer. The symptom is the incident review quoted above, repeated for months.
  • A service re-decides mid-trace, producing partial traces that mislead rather than inform.
  • Percentage sampling under a traffic spike multiplies telemetry cost exactly when the system is under stress, which is when the telemetry pipeline can least afford it.
  • Counts computed from sampled traces without the rate recorded. Teams reason for months about "how often this happens" from a number that is off by a factor nobody remembers setting.
  • A debug override with no authentication becomes a way for any caller to force full tracing, and therefore a cost-amplification vector.

Trade-offs

Choose Gains Pays
Head-based No buffering tier, cost saved at the source, trivial to operate Cannot select on outcome; rare failures are lost
Tail-based Keeps every error and slow request within the same budget Stateful collector tier, trace-ID-affine routing, memory sized on p99 duration

When not to use it

Drop it when the failures you need are discovered late: a request that returns HTTP 200 and is wrong, a slow hop three services deep that the entry point cannot see, an error raised in an async continuation. If the incident reviews keep saying the trace was missing, the rate is not the problem and raising it will not help — the decision point is the problem, and only tail-based sampling moves it.

Also avoid it as the only mechanism in a system with pronounced request-cost skew. If 0.1% of requests do 60% of the work, uniform head sampling will almost never capture one, and the expensive path stays invisible. A per-route rate fixes that cheaply, provided the route is known at the head.

Conversely, do not abandon it on principle for a mesh of fifteen services one team owns. The buffer tier, the hash-routing policy and the eviction timeout are real operational surface, and "ask the team that owns the slow hop" is still a viable substitute at that size.

Interview question

Q: Your tracing bill is fine and your coverage of incidents is terrible. The team proposes raising the sample rate 20x. What do you say?

What a strong answer covers: that uniform sampling is proportional, so 20x the cost buys 20x the chance of a coincidence rather than the trace you need · that the decision point, not the rate, is the constraint · the per-route and error-at-the-root refinements that are nearly free · what tail-based sampling would cost in buffer memory sized on p99 trace duration and in hash-affine routing · and the debug-override header as the cheapest partial answer available this week.

Quick check

Quiz: Why is raising uniform head-based sampling from 0.1% to 2% an ineffective fix for missing traces of rare failures? — Because it keeps 2% of failures as well as 2% of successes; the ratio is unchanged and rare events stay rare.

Flashcard: Where is a head-based sampling decision made, and what must happen to it afterwards? — At the root span, before the outcome exists; it must propagate unchanged in the trace context, because a service that re-decides produces partial traces.