Structured Logging
Machine-parseable events with stable names and consistent fields.
3 to work through
-
intermediate
A platform moves from free-text logs to structured logs. What becomes possible, and what discipline must accompany it to avoid making things worse?
2 min answer -
intermediate
A ride-hailing platform of Grab's shape standardises on structured logs. After a release, the on-call dashboard that counts failed trips by city reads zero for three services and normal numbers for the rest. No errors are being reported anywhere. What broke, and what should have caught it?
2 min answer -
advanced
At 09:40 a release begins logging the whole inbound request as nested JSON, including a customer-supplied metadata map. At 10:05 the log store starts rejecting documents for that index with an illegal-argument error. By 10:20 no service writing to that index can log at all, including services that changed nothing. What happened, and which design decision allowed it?
2 min answer
3 terms in this topic
Log Schema Consistency
Enforcing the same field names, types and semantics for structured log records across every service, so cross-service queries are possible.
conceptLog Schema Drift
The silent breakage of dashboards and alerts when a service changes the name, type or nesting of a structured log field, producing empty results rath…
practiceStructured Logging
Emitting logs as machine-parseable key-value records rather than formatted prose, so they can be queried, aggregated and correlated.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.