Log Sampling Budget
also called Log Volume Budget, Log Ingest Quota
An owned per-service allowance of log volume, spent by deciding which record classes are kept whole and which are sampled by trace - the control that turns an invisible shared cost into a decision a team makes deliberately.
A platform's log bill grows faster than its traffic for four quarters. Nobody added a logging feature; twelve teams each added a few fields and a few lines, every change too small to argue about. The platform team's only lever is global: raise the sampling rate for everyone, or cut retention for everyone. Both are decisions taken by the wrong people.
A log sampling budget replaces the global lever with a per-service allowance plus rules on how it may be spent. The allowance is a number, in gigabytes per day or a share of ingest, owned by a team. The rules name which classes are exempt and how sampling must be implemented, because a budget without them gets met in ways that destroy the logs while still paying for them.
Why it matters
Log volume is an externality: the cost lands on a shared bill and the benefit lands on the team that added the line, so the rational individual choice is always to log more. Any control that does not attribute cost to a team fails for that reason. Both usual global controls do damage — uniform sampling destroys reconstructability, uniform retention cuts obligations along with noise.
The lever is also large and safe: in most estates 70 to 90% of volume is successful-request INFO, and the error records investigations need are a rounding error on the total.
Implementation patterns
- Exempt classes first, rate second. Errors, audit and access records, and anything with a retention obligation stay at 100% and outside the budget. Choosing a rate before naming exemptions is how teams end up sampling their own error logs.
- Sample by trace, never by line. Hash the trace id against a threshold so every service makes the same decision for one request. OpenTelemetry's consistent probability sampling gives that threshold a wire format in the
thfield oftracestate, so the decision travels and counts can be scaled back up. - Record the rate on the survivors, because a sampled log with no rate cannot be turned back into a count and someone will try.
- Attribute volume per service in a dashboard the owning team sees; attribution changes behaviour faster than any limit.
Industry example
The economics are visible from the supply side. Datadog's Husky posts in 2022 describe an event store built so that retention is a pricing decision rather than a capacity project, with columnar files on object storage and continuous compaction to keep scans efficient. A budget addresses the same problem from the demand side: someone pays per byte per day, and the question is whether whoever chooses the bytes sees the bill.
Failure scenarios
- Per-line random sampling. The bill falls 90% and every investigation now works from a tenth of each request's records, so nothing can be reconstructed. The budget was met and the capability destroyed.
- Sampled errors. A uniform rate over a 0.1% error class leaves roughly one error record per 10000 requests, which cannot support a customer report.
- A debug toggle outside the budget. One engineer flips a service to DEBUG, volume rises 20 to 50 times, and the shared quota is gone in minutes.
Trade-offs
The budget buys a predictable bill, tenant isolation and a conversation about value that would otherwise not happen. It pays in three currencies: every investigation carries a caveat about what was kept, the platform team owns an allocation process that is political work, and the machinery can be misconfigured with a failure mode of missing data rather than an error. The flip condition is whether logging is in your top few cost lines, because below that a budget adds a permanent caveat to every future investigation for a saving nobody will notice.
When not to use it
A small estate with a modest bill should not build this, and neither should a team whose logs are predominantly audit or transaction evidence, because almost everything is exempt. Never apply a cap before per-service attribution exists, because a cap without attribution moves the pain to whichever team next adds a useful log line. Attribution alone, published where owners see it, often removes enough volume that the cap is never needed.
Interview question
Q: Your log bill must fall 60% this quarter across 40 teams, without losing the ability to diagnose incidents. What do you do in what order, and what do you refuse?
What a strong answer covers: attribution before enforcement, per service and per log statement; naming exempt classes before choosing any rate; trace-deterministic sampling of successful-path INFO with the rate recorded on survivors; retention tiering so days 4 to 90 live in object storage rather than a search index; and refusing to sample errors, refusing to cut retention on obligated records, and refusing a global rate change as the first move.
Quick check
Quiz: Why is sampling logs at 10% per line worse than keeping 10% of traces whole? Per-line sampling leaves a tenth of every request's records, so nothing can be reconstructed while you still pay a tenth of the bill; per-trace sampling leaves whole conversations for the same money.
Flashcard: Which log classes are exempt from a sampling budget? — Errors, audit and access records, and anything with a retention obligation. They are a rounding error on volume and the reason the logs exist.