Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
4 to work through
-
intermediate Multiple choice
A design platform of the kind Canva runs exports 40 metrics per service. An engineer adds a `pod_name` label so a noisy pod can be identified. The service runs 600 pods and deploys twice a day, so pod names turn over completely every 12 hours. Retention is 30 days. Roughly how many distinct series does that one label create over the retention window?
2 min answer -
intermediate
A platform's log volume has grown until log storage is one of its largest infrastructure costs. What should change, and what should not?
2 min answer -
advanced
A platform's observability spend has grown to a substantial fraction of its infrastructure budget with no single team responsible. What are the drivers, and how should this be controlled without losing visibility?
2 min answer -
advanced
An observability bill has grown to a significant fraction of infrastructure spend. Diagnose it systematically and describe the reductions that do not lose diagnostic power.
3 min answer
2 terms in this topic
Telemetry Cost Control
Managing the spend on logs, metrics and traces, which in mature estates can approach or exceed the cost of the infrastructure being observed.
practiceTelemetry Cost Management
Controlling observability spend through sampling, retention tiering and cardinality limits without losing diagnostic capability.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.