Logging
What to log, at what level, and what must never appear in a log.
4 to work through
-
beginner
A service emits one INFO log line per request and increments one counter per request. At 3,000 requests per second the two describe exactly the same traffic. Why does keeping the log cost orders of magnitude more than keeping the counter, and when is the log still the right thing to emit?
3 min answer -
intermediate
A discussion platform's log volume grows faster than its traffic and now costs more than the compute generating it. What is driving the growth, and what should change?
2 min answer -
intermediate
A platform team stops shipping logs from inside each application process to the log vendor and instead writes JSON to stdout for a node agent to collect. Crash-time logs now survive and the request path no longer touches the network. What has the team given up, and when does that bill arrive?
2 min answer -
intermediate
Your logging bill has tripled alongside traffic growth. Name four changes in order of effectiveness.
2 min answer
3 terms in this topic
Log Level Discipline
Consistent semantics for log severity so that levels can be used for routing, alerting and cost control.
practiceLog Sampling Budget
An owned per-service allowance of log volume, spent by deciding which record classes are kept whole and which are sampled by trace - the control that…
practiceLogging
Emitting events a human or a query can reason about later — where the discipline is structure, sampling and what you refuse to log.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.