practice

Structured Logging

Emitting logs as machine-parseable key-value records rather than formatted prose, so they can be queried, aggregated and correlated.

loggingobservabilityoperations

The difference is between a log line that reads well to a person reading one line, and a log event that can be queried across a billion of them. At any real scale only the second is useful, because nobody reads logs sequentially — they filter, aggregate and correlate.

A structured event carries the fields as data: timestamp, level, service, version, trace and span identifiers, user or tenant where appropriate, the operation, the outcome, the duration, and any domain identifiers relevant to the event. The message becomes one field among many rather than the container that everything must be parsed out of.

What this enables is the difference between a five-minute investigation and an afternoon: filtering to one tenant's failures, grouping error counts by version to identify a bad deployment, computing the latency distribution of an operation from the logs themselves, and joining directly to traces via the identifier.

Two disciplines make it hold. A shared schema across services, since the query value collapses if every team names the same concept differently — the correlation identifier in particular must be identical everywhere. And deliberate exclusion of sensitive data, because structured logs are easier to search, which cuts both ways: personal data in logs is now indexed, retained, and replicated to wherever logs go, which is frequently a third-party service in another jurisdiction.