OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.
4 to work through
-
intermediate
A platform with services in five languages and three monitoring vendors considers adopting OpenTelemetry. What does it solve, and what is the realistic migration cost?
2 min answer -
intermediate
An organisation with several existing telemetry systems considers adopting OpenTelemetry. What does it actually solve, and what does adopting it not fix?
2 min answer -
advanced
A fleet exports OTLP to a gateway collector deployment with memory_limiter first in the pipeline, then a batch processor, then an exporter with a sending queue. The telemetry backend starts answering in 8 seconds instead of 80 ms, and request volume doubles at the same time. What happens second by second, and what stops it?
2 min answer -
advanced
A platform team standardises every service on OTLP push to a collector and switches off Prometheus scraping. Metrics still arrive and dashboards still work. What has the team given up, and when does that bill arrive?
3 min answer
3 terms in this topic
Exemplar
A trace identifier attached to a metric data point, linking an aggregate measurement directly to a concrete request that produced it.
conceptMetric Temporality
Whether a reported metric point carries a running total since a start time or only the change during one interval - the choice that decides whether p…
toolOpenTelemetry
A vendor-neutral standard and toolset for generating, collecting and exporting traces, metrics and logs.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.