Trace Exemplar
also called Metric Exemplar, Sampled Trace Pointer
A trace ID attached to a histogram bucket or counter sample, letting a responder jump from an aggregate graph straight into one real request that produced it - without adding a single active series.
The oldest gap in a metrics dashboard: the p99 latency panel shows a spike, and there is no way to get from the spike to a request. The graph proves something happened and contains no example of it. The responder's next move is to guess a time range and go searching in logs or traces, which is the slow part of most incidents.
The instinct is to add labels until the metric can answer the question. That is the move that destroys a metrics platform, because label cardinality is multiplicative and series are expensive.
An exemplar is the other answer. When a measurement is recorded, the instrumentation attaches the trace ID of that specific request to the bucket it landed in. The series is unchanged — same labels, same cost — but the bucket now carries a handful of pointers to real requests. Click the spike, land in a trace.
Why it matters
It breaks the trade-off between aggregate visibility and individual detail at zero cardinality cost. A customer_id label on a histogram with 288,000 existing combinations and 20,000 customers is tens of billions of series. An exemplar is a pointer stored alongside a sample: a handful of bytes per bucket per interval, and no new series at all.
It also fixes the direction of the workflow. Traces are usually entered by searching — a service name, a time range, a guess at a filter. Exemplars let the responder enter from the symptom: this bucket, the slow one, right now. The request you land on is guaranteed to be an instance of the problem, which searching never guarantees.
Implementation patterns
- Attach at the slow buckets. Most implementations keep one or a few exemplars per bucket per interval. The ones that matter are in the high-latency and error buckets, so prefer a policy that always retains an exemplar there over a uniform one.
- Only sampled traces are reachable. An exemplar pointing at a trace that was dropped at the head is a dead link. Either coordinate the policies — keep the trace if its measurement landed in an outlier bucket — or accept that exemplars work for the sampled fraction.
- Propagate the trace ID into logs too, so the same click reaches logs and traces alike.
- Honour the standards. OpenTelemetry defines exemplars on metric data points, and Prometheus supports them on the OpenMetrics exposition format. Using the standard means Grafana and similar tools render the click-through without custom work.
- Budget retention separately. Exemplars are cheap per sample and are retained with the metric, which is typically far longer than traces. A pointer that outlives its trace is a dead link, so align the two or label the metric's drill-down window.
Industry example
Exemplars exist because of the cost described in eBay's published account of Sherlock.io (2022): a tier holding roughly 3 billion active series fed by about 40 million samples per second. At that scale, "add a label so we can drill down" is not a decision with a cost, it is a decision with no possible implementation. The OpenTelemetry and OpenMetrics specifications added exemplars to give the drill-down back without the series, and that is the correct way to read the feature: not as a convenience, but as the structural answer to high-cardinality drill-down on a metrics system that cannot afford high cardinality.
Failure scenarios
- Dead links at scale. Exemplars reference traces that sampling discarded, so most clicks go nowhere and the team stops clicking. This is the single most common reason exemplars are judged useless.
- Unrepresentative examples. One exemplar per bucket per interval is one request, not a distribution. A responder who generalises from it diagnoses the wrong cause — the exemplar is a starting point, never evidence of prevalence.
- Quiet cardinality through the back door. Implementations that attach exemplar labels rather than a bare trace ID can reintroduce the cost the mechanism exists to avoid.
- Retention mismatch, where metrics are kept for a year and traces for a week, so every exemplar older than seven days is a broken link and nobody has said so.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Exemplars | Drill-down at zero series cost; entry from the symptom | Only sampled traces reachable; one example, not a distribution |
| High-cardinality labels | Query any dimension, count accurately | Multiplicative series cost; degrades queries for every tenant |
| Wide events | Any field queryable at read time, cardinality free | A second storage system, shorter retention, higher per-byte cost |
When not to use it
If you already run wide events per request, exemplars add little. The event carries the identifiers as columns, so the drill-down is a query rather than a pointer, and it works for every request rather than the sampled few. Running both is not wrong; investing in exemplar plumbing when the event store answers the same question better is.
They are also the wrong answer when the question is about prevalence rather than instance. "How many customers are affected, and which?" is a counting question, and a pointer to one request cannot answer it. That needs an aggregated dimension — a narrow per-tenant metric for the tenants you owe an SLA, or the event store. Reaching for an exemplar there produces a confident answer from a sample of one.
Interview question
Q: A team wants per-customer latency drill-down and your metrics tier cannot take the cardinality. What do you offer, and what does each option fail to do?
What a strong answer covers: exemplars as the zero-cardinality drill-down, with the honest caveat that only sampled traces are reachable and one exemplar is not a distribution · wide events as the structurally correct home for per-request dimensions, at the cost of a second system · a narrow tiered metric for the small set of customers under contract, which is usually the real requirement · and the recognition that the three answer different questions — instance, arbitrary query, and count — so the right move is to establish which one was being asked.
Quick check
Quiz: What does attaching an exemplar to a histogram bucket cost in active series? — Nothing. It is a pointer stored with the sample, not a new label combination.
Flashcard: Why do exemplars so often produce dead links? — They point at traces the sampling policy discarded. Coordinate the policies — retain the trace when its measurement lands in an outlier bucket — or accept coverage limited to the sampled fraction.