Metrics
Counters, gauges and histograms, and percentiles rather than means.
4 to work through
-
intermediate
A live platform's dashboards show mean latency, which stays flat during an incident where many users experience severe delays. Why do averages hide this, and what should be measured?
2 min answer -
intermediate
A multi-tenant SaaS platform cannot tell which customer caused a performance problem. What must be present in the telemetry, and what does that cost?
2 min answer -
intermediate
Etsy released StatsD in 2011: a small daemon that receives metric samples from applications over UDP, aggregates them in memory, and flushes to the metrics backend on a fixed interval. Why was UDP the right choice for that design, and what did it make permanently impossible?
3 min answer -
intermediate
Every service dashboard shows healthy metrics while users cannot complete a purchase. How is that possible and what would have detected it?
2 min answer
4 terms in this topic
Counter Reset Handling
The rule that a decrease in a cumulative counter is interpreted as a process restart rather than as negative work, which is what allows rates to surv…
conceptCounter, Gauge and Histogram
The three fundamental metric types, distinguished by what they represent over time and by which aggregations are valid.
conceptMetric Staleness
The rules deciding how long a time series keeps answering queries after it stops being written - which determines whether a dead process's last value…
conceptMetrics
Cheap pre-aggregated numeric time series — excellent for knowing something is wrong, structurally unable to tell you which request.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.