A team moves from human-readable log lines to structured JSON logs, with every line carrying service, version, request ID, user ID and a set of typed fields. Querying gets much better. What have they given up, and when does that bill arrive?
Show the full answer Hide the answer
What is gained, concretely
A log line becomes a record you can query rather than text you have to match. status >= 500 AND region = "eu-west-1" AND version = "4.21.0" is a query that returns an exact answer in seconds. The equivalent against free-text lines is a regular expression that is wrong in a way you discover later.
The second gain is bigger and less discussed: the fields are consistent across services, so a query written against one service works against all of them. That is what makes a central log platform useful rather than a pile of nine different formats.
What is paid
Volume, by a factor of roughly 2 to 5. A 90-character log line becomes a 400–600 byte JSON object once every line repeats the service name, the version, the host, the trace ID and the field keys. The keys are the surprise: {"request_id": "...", "user_id": "...", "duration_ms": 42} spends a meaningful fraction of its bytes on the words request_id and user_id, re-sent on every line. At 50,000 lines per second, a 3x expansion is roughly an extra 50 MB/s to ship, index and store — and log platforms generally charge on ingested volume, so this lands directly on the bill.
Human readability during a crisis. When you are tailing a log on a host because the aggregation layer is itself broken, JSON is genuinely harder to read than a sentence. The mitigation is real and is a mitigation: emit JSON in production and pretty-printed text in local development, and keep a jq incantation in the runbook.
Schema discipline, permanently. The value of structure comes from fields meaning the same thing everywhere. The moment one service logs userId, another user_id and a third uid, you have the old problem with more syntax. This requires a shared library and a review habit, which is organisational cost rather than technical cost, and it is the part that decays.
When the cost becomes visible
Not at adoption, which is the trap. It arrives at the first bill after the first traffic doubling, because log cost scales with request volume while the perceived benefit does not. Teams typically discover it as a line item that has quietly become comparable to their compute spend.
It also arrives during the first incident where the logging pipeline is saturated. Three times the bytes means the buffer that held 90 seconds of logs now holds 30, so during the incident that produces a burst of error logging, the logs are dropped exactly when they are needed.
When not to do this, and how to keep the cost down
For a single service with one developer, plain lines in a file are fine and JSON is ceremony: the query advantage only appears once you are searching across services you did not write. The threshold is the second service and the first shared log platform. Below it, the cost is real and the benefit is theoretical.
Above it you will not reverse the decision, and these are the levers that make the cost manageable:
- Sample the routine lines, keep all the errors. Successful requests at one in a hundred, errors at one in one. This is where the volume is, and it is the single largest saving available.
- Log events, not steps. One wide line per request with twenty fields beats twenty lines with one field each: it costs less, and it is more queryable because the fields are correlated on one record.
- Short keys, or a schema the platform knows.
svcinstead ofservice_nameacross billions of lines is not micro-optimisation at this volume. - Tier retention. Full fidelity for 7 days, errors for 30, aggregated counts for a year. Almost nobody queries raw logs older than a week, and almost everybody pays to store them.
The decision rule: structured logging is correct for anything more than one service, and the volume levers above are not optional extras — they are part of the design. A team that adopts the format without adopting the sampling has bought the cost and deferred the discipline.