Correlation IDs
One identifier propagated through every hop and every log line.
3 to work through
-
beginner Multiple choice
A customer reports an error at 14:32 yesterday. What must have been built for the investigation to take two minutes rather than two hours?
2 min answer -
intermediate
A logistics platform's request crosses synchronous services, message queues, scheduled batch jobs and third-party callbacks. Tracing works within services and breaks between them. What is missing?
2 min answer -
intermediate
An API platform's requests trigger asynchronous work, webhook deliveries and retries, sometimes hours later. How should correlation identifiers be designed so a customer question can be answered end to end?
2 min answer
3 terms in this topic
Business Correlation Identifier
A domain identifier - shipment, order, trip, claim - carried on every log line, event, job and external call, answering the question distributed trac…
patternContext Propagation
Carrying request-scoped identifiers and metadata across every service, thread and asynchronous boundary so a single flow remains traceable end to end.
patternCorrelation IDs
A single identifier attached at the edge and carried through every hop, synchronous and asynchronous, that stitches an investigation together.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.