A request fails somewhere in a chain of five services. Each service logs the error clearly, with a timestamp and a message. Why is finding the five lines that belong to that one request still hard, and what single change fixes it?
Show the full answer Hide the answer
The mechanism
Logs are written per-process, and a request is a path across processes. Each of the five services wrote a true, well-formed line about what it saw. Nothing connects them.
The instinct is to use time. It does not work, for a reason worth understanding: at even modest traffic the services are handling many requests concurrently, so the window around a timestamp contains hundreds of other requests' lines, and the five you want are interleaved with them. Worse, the clocks on five hosts are not identical. NTP keeps them within milliseconds of each other on a good day, which is far coarser than the gaps between the lines you are trying to order. Sorting by timestamp across hosts can produce an effect that appears to precede its cause.
So the reality is: you have the information, it is written down, and you cannot assemble it. The incident is spent grepping.
The change
Generate one identifier at the edge and pass it to every service in the chain, and make every log line carry it.
The identifier is created once, at the first point the request enters your system — the load balancer, the API gateway, the front-end service. Each service reads it from the incoming request headers, attaches it to everything it logs, and puts it on the outgoing headers of every call it makes. Then correlation_id=7f3a… in the log search returns exactly the five lines, in order, across all five services.
The load-bearing word is propagate. A service that generates its own ID instead of reading the incoming one breaks the chain at that point, and this is the single most common way the mechanism is implemented and still fails.
The consequence people miss
It has to cover the asynchronous hops too. If service C publishes a message to a queue and service D consumes it a second later, the ID must travel in the message metadata, or the trail ends at the queue — which is usually where the interesting failures live. Likewise for retries, scheduled jobs kicked off by a request, and calls to external partners whose logs you will later be comparing against.
The second thing people miss: put it in the error the user sees. A support ticket that says "something went wrong, reference 7f3a9c" is a one-query investigation. A ticket that says "it was broken this morning" is an afternoon.
When not to build this by hand
If you already run distributed tracing, you largely have this: the trace ID serves as the correlation ID and the tracing library propagates it through the same headers, so writing your own threading of an ID through every call is duplicated work that will drift. Adopt the W3C Trace Context standard — a traceparent header, standardised in 2020 — rather than inventing X-My-Company-Request-Id, because the standard header is the one your proxies, service mesh and vendor agents already understand, and a custom header is one that every new component has to be taught.
The reason to understand correlation IDs as a separate idea is that tracing is usually sampled and logs usually are not. A trace retained at 2% leaves 98% of requests with no trace at all, and for those the ID in the log line is the only thread back through the system. So the decision rule is: log the trace ID on every line whether or not that trace was sampled, and the two mechanisms cost you one field.
What it costs is small and not zero: a few bytes on every log line and every outbound request header, and the discipline that every new service and every new transport propagates it. The discipline is the real price — the mechanism is trivial and it fails the moment one service in the chain forgets, which is why it belongs in a shared middleware library rather than in each service's code.
Common weak answers
- "Search by user ID." Better than timestamps, and a user makes many requests. It narrows the haystack; it does not identify the request.
- "Add more logging." More lines, same problem, now slower to search and more expensive to store.
- "Use the load balancer's request ID." A reasonable source for the value, and it only helps if every downstream service propagates it. The identifier is not the hard part; the propagation discipline is.