advanced 3 min answer

To stop losing the traces that matter, a team moves from 5% head sampling to capturing every span and deciding in its own collector tier, keeping 2% after the decision. Vendor ingest falls by 60%. What has the team given up, and when does that bill arrive?

tracingtail-samplingcollectorscross-zonetelemetry-cost
Show the full answer Hide the answer

What is gained, quantified

Head sampling decides before the outcome is known, so 5% of errors survive and 95% of the interesting traces are destroyed at birth. Tail sampling decides after the trace completes, so the 2% kept can be every error, every trace over the latency threshold, and a thin slice of the boring remainder. Diagnostic value per stored byte goes up by an order of magnitude, and the vendor line falls 60% at the same time. Both of those are real.

What is paid

A tail decision needs the whole trace, so every span is buffered until the trace finishes or a timeout fires, and all spans of one trace must arrive at the same collector instance because the decision cannot be made from a fragment. That turns a stateless exporter into a stateful, routed fleet, and three costs follow.

  • Memory. 10000 requests a second at 20 spans of roughly 500 bytes is about 100 MB a second of span data, so a 30-second decision window holds on the order of 3 GB in flight before headroom, replication or head-of-line stalls.
  • Network. Routing spans by trace id adds a hop. Cross-zone transfer is billed in both directions, so a routing layer that crosses availability zones doubles the transfer you modelled rather than adding to it. Keep the routing hop in-zone and the meter never starts.
  • Operations. You now run a tier whose failure silently drops telemetry, during the incident when it is the only evidence.

When the cost becomes visible

Not on a quiet Tuesday. It arrives during the incident the investment was made for. Load rises, traces lengthen, and traces that now exceed the decision timeout are evicted incomplete, so the sampler keeps less of exactly the pathology you were buying. At the same time memory pressure in the collector tier rises with trace duration, not with request rate, so the fleet sized against normal traffic is undersized against slow traffic. The second-order version is worse: the collector is scheduled on the same nodes as the workload, so a node under pressure loses its evidence first.

How to keep the option to reverse

  • Keep unsampled aggregates. Request counts, error counts and latency histograms derived at the edge, independent of the trace pipeline, so a collector failure degrades to aggregate signals rather than to nothing.
  • Record the effective sample rate on every stored span, so counts can be reweighted. Without it every number computed from traces is wrong by an unknown factor.
  • Set the decision timeout from the measured p99.9 trace duration, and alert on evictions. Evictions are the signal that the pipeline is lying to you.
  • Pin collectors in-zone and make the routing layer zone-local, so the transfer meter is a design choice rather than a surprise.

When this is the wrong answer

If capturing everything at the vendor costs less than about 10% of your infrastructure bill, pay the vendor and skip the fleet. A collector tier is roughly a quarter to a half of an engineer in perpetuity plus its own compute; below a few terabytes of spans a day that exceeds the saving. Prefer the intermediate move instead, which is error-and-latency-biased head sampling: keep every error and every slow request, sample successes hard. It captures most of the diagnostic gain with a stateless pipeline and no buffering at all, which is why it is the step to try in production before committing to a routed collector tier.