Span Link
also called Trace Link, Span Reference
A pointer from one span to another span in a different trace, recording a causal relationship where no single enclosing parent exists - the mechanism that keeps batch jobs, fan-in consumers and delayed asynchronous work connected without fusing unrelated traces into one.
A nightly reconciliation job reads 5000 orders that each arrived hours earlier with their own trace. If parent-child is the only relationship available, both options are bad. Parent the job to one order and that order's trace swallows the whole batch, a waterfall nobody can read. Leave the job unparented and it is an island: when it marks an order wrong, there is no path from the order's trace to the run that did it.
A span link is the third option: a typed pointer carrying another span's trace id and span id plus optional attributes, with no claim that either encloses the other. The OpenTelemetry specification states that a parent represents a single enclosing scope, which is why it advises against parenting in batch and scatter-gather shapes.
Why it matters
Queues, batches, retries and fan-in consumers are all points where a trace either ends or swallows something it should not, and the choice made there decides whether an investigation can cross the boundary at all.
Getting it wrong does more than lose a link. A batch parented into one request's trace inflates that request's duration to hours, which poisons every latency statistic derived from trace data and makes the trace store's own percentile views untrustworthy.
Implementation patterns
- Queue producer and consumer. The consumer starts a new trace and links back to the producer. Parent instead only when the producer actually waits for the consumer.
- Batch and fan-in. Link to each contributing span; above a few hundred, link to a summary span and record the full id list as data.
- Retry chains. A retry links to the attempt it replaces, so "how many times did this run" is answerable without inferring it from timestamps.
- Relationship attributes. A bare link says only "related". Name the edge — batch member, retry of, triggered by — so queries can distinguish kinds.
Industry example
The anchor is a specification rather than a company blog, which suits a primitive: OpenTelemetry's messaging semantic conventions describe when a consumer should link rather than parent, and the guidance dates from the work that stabilised the tracing API in 2021. A food-delivery platform in Zomato's mould shows the stakes in production. An order-status event is produced by the order service, delivered by a webhook fleet, retried against a tablet, and later swept by reconciliation: four systems, four traces, and one customer complaint that needs all four.
Failure scenarios
- Dangling links. A link into an unsampled trace resolves to nothing, so the engineer clicks through to an empty page. This is the strongest operational case for consistent sampling across the estate.
- Link explosion. A span with 5000 links costs one lookup per link to resolve and most user interfaces will refuse to render it usefully.
- Links as a substitute for a domain identifier. Trace data is sampled and often kept 7 days; a question about an order over three days needs an order id on durable records.
- Silent backend gaps. Some vendors accept links on ingest and never surface them, so the instrumentation buys nothing until the rendering is verified.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Link (separate traces) | Readable traces; independent sampling; honest durations | An extra hop for the investigator; dangling links when sampling diverges |
| Parent (one trace) | One waterfall shows everything | Hour-long traces; distorted percentiles; one arbitrary parent among many |
| Domain correlation id | Survives sampling and retention; spans days | No timing detail; discipline on every record |
When not to use it
For a synchronous call where one operation genuinely waits for another, a parent is correct and a link is strictly worse, because you lose the enclosing duration that makes a waterfall interpretable. For questions spanning days, or that must survive sampling, prefer a domain identifier on every log line and event. And if the backend does not render links, prove the rendering on one pipeline before asking twelve teams to instrument.
Interview question
Q: A consumer reads batches of 500 messages from a queue and processes them in one transaction. Producers are 12 different services. Describe the trace structure you would instrument, and what a reader loses under your choice.
What a strong answer covers: a new trace per batch, parented to nothing, with one link per message carrying a relationship attribute; that 500 links is near the practical limit, so linking to a summary and recording ids as data is the scaling move; that the reader loses a single view of one message's life and gains traces whose durations mean something; that sampling must be consistent or the links dangle; and that a message id on every log line is what answers "what happened to this message" days later.
Quick check
Quiz: Why does the OpenTelemetry specification advise against setting a parent in batch and scatter-gather scenarios? A parent asserts one span that encloses the child, which is false when there are many originating spans that finished long before; a link carries the causal edge without the false enclosure.
Flashcard: A nightly job processes 5000 orders that each had their own trace. Parent or link? — Link. Parenting picks one arbitrary parent and swallows the batch into that order's trace; a link into an unsampled trace is the failure mode to guard.