intermediate 2 min answer

A platform team stops shipping logs from inside each application process to the log vendor and instead writes JSON to stdout for a node agent to collect. Crash-time logs now survive and the request path no longer touches the network. What has the team given up, and when does that bill arrive?

loggingkuberneteslog-rotationtruncationbackpressure
Show the full answer Hide the answer

What is gained, quantified

Writing to stdout is a write to a local pipe, so a log line emitted microseconds before an out-of-memory kill is already out of the process. The old design lost exactly those lines, which are the ones that explain the crash. Connection count drops from one vendor client per pod to one agent per node, typically a 20:1 to 100:1 reduction, and no application thread ever blocks on a vendor's slow response.

What is paid

Long records are split by the container runtime. A log line longer than about 16 KB is broken into chunks; the CRI log format tags every chunk but the last as partial (P) and the final one as full (F). If the collector is not configured to reconcile them, a 60 KB stack trace arrives as four fragments, each of which fails JSON parsing and is dropped. The symptom is that logs are silently missing for precisely the largest errors, which is the opposite of what anyone would guess.

The node has a fixed log budget and enforces it by deletion. The kubelet rotates container logs at a default 10 MiB per file keeping 5 files, and the rotation routine runs about every 10 seconds. A pod emitting 50 MB/s during a retry storm rolls through all 50 MiB in a second, so anything the agent had not yet read is gone. The loss happens during the incident, which is the only time the logs mattered.

Per-record backpressure disappears. The application's write succeeds whatever happens downstream, so the app cannot know its logs were dropped, and the pipeline's drop counters live in the agent rather than in the service that lost the data.

When the bill arrives

Not at migration. It arrives at the first incident with a log burst or a large payload: a retry storm, a deserialisation failure that dumps request bodies, a dependency returning verbose errors. The bill is denominated in the exact records the postmortem needs.

What to do about each

Cap a single record in the logging library — truncate the payload at roughly 8 KB and put the full object in an object store with a pointer field. Raise containerLogMaxSize for the few pods that legitimately burst, and treat a pod that can fill 50 MiB in a second as a logging-volume defect rather than a rotation-settings problem. Configure the collector's multiline or partial-message mode for the runtime in use, and alert on the agent's parse-failure and dropped-record counters, because those are the only signals that this failure mode exists.

When this is the wrong trade

A handful of low-volume services on long-lived hosts, shipping to a sink with a local disk buffer, are better served by direct shipping: the library gets a delivery acknowledgement, retries with its own buffer, and no runtime sits between the record and the pipeline. The stdout pattern earns its costs when pod count is large and lifetimes are short, because then the per-process client and its buffer are the bigger problem. Below roughly a few dozen long-running processes, it is machinery bought for a scale that does not exist.