Distributed Tracing
Reconstructing one request's path across every service it touched.
4 to work through
-
intermediate
A request takes 3 seconds. Every individual service reports healthy latency. How does tracing resolve this and what must have been instrumented?
2 min answer -
advanced
A marketplace adopts distributed tracing but engineers rarely use it during incidents. What typically causes low adoption, and what makes tracing actually useful?
2 min answer -
advanced
A team enabling distributed tracing must decide between head-based and tail-based sampling. Compare them, and explain what makes traces useful beyond a single request's timeline.
3 min answer -
advanced
You are introducing distributed tracing across 40 services owned by 12 teams. Plan the adoption.
2 min answer
5 terms in this topic
Critical Path Analysis
Identifying which spans in a trace actually block the response, as opposed to running in parallel, so optimisation effort targets the work that deter…
conceptDistributed Trace
A causally-linked record of one request's path across services, composed of spans that carry timing, attributes and parent relationships.
toolDistributed Tracing in Practice
Following one request across every service it touches — the only tool that answers "where did the time go" in a distributed system.
practiceSpan Attribute Budget
A deliberate limit on how many spans a request emits and how many attributes each carries, set so that traces stay readable and affordable rather tha…
conceptSpan Link
A pointer from one span to another span in a different trace, recording a causal relationship where no single enclosing parent exists - the mechanism…
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.