Sampling
Head-based versus tail-based, and keeping the traces that matter.
3 to work through
-
advanced
A commerce API serves 2000 requests per second and keeps 1% of traces. The team wants to watch a 0.5% checkout error rate from the trace store and page when it moves by a tenth of itself. Roughly how much traffic must the sample cover before that reading is trustworthy, and what does the answer say about the alert?
3 min answer -
advanced
A platform with extreme traffic spikes needs distributed tracing. Head-based sampling loses the interesting traces and 100% sampling is unaffordable. What strategy resolves this?
2 min answer -
advanced Multiple choice
An edge platform handles enormous request volume and cannot retain telemetry for every request. Which sampling strategy should it use, and what breaks under naive uniform sampling?
2 min answer
2 terms in this topic
Tail-Based Sampling
Deciding whether to keep a trace after it completes, so that errors and slow requests are always retained while ordinary traffic is sampled cheaply.
practiceTrace Sampling Strategy
Deciding which traces to keep, since retaining all of them at scale is unaffordable and retaining a random few loses exactly the interesting ones.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.