A platform injects a log-shipping sidecar with a 128 MiB limit into every pod. During an evening peak the sidecar's in-memory buffer grows faster than it can flush, the pod exceeds its limit, the kubelet kills the container and a product team's checkout service drops requests for eleven minutes. Whose reliability budget is spent and what does that imply about how the platform ships that sidecar?
Show the full answer Hide the answer
What is being tested
Whether you can see that a platform stops being a service next to the product the moment its code runs inside the product's failure domain. The product team's error budget is spent either way — their users saw the failure — but the decision that caused it was the platform's, and the platform is the only party that can fix it for everyone.
The reasoning
An injected sidecar makes the platform a runtime dependency of every service, so the platform inherits the product's pager and the product inherits the platform's bugs. Three consequences follow, and they are what the answer is really about:
- The failure is fleet-wide by construction. The same default limit and the same buffering code are in every pod. What happened to checkout at peak will happen to every service whose log volume crosses the same threshold, in log-volume order. One team noticed; the defect is estate-wide.
- The blast radius is set by the default, not by the incident. A 128 MiB limit and an unbounded buffer is a design choice made once and multiplied by the fleet. At 2000 pods that default also reserves roughly 250 GB of cluster memory, which is the cost side of the same decision.
- The rollout mechanism is the control. Sidecar images are usually pinned by the injector, so a bad version reaches pods as they restart rather than all at once. That is an advantage only if the platform can see version spread per namespace and stop it; otherwise it is a slow, invisible rollout.
The practical rule: for anything the platform places inside someone else's process or pod, the platform contract has to name the pager. Write down which symptoms page the platform (the sidecar is killed, cannot flush, or adds latency), which page the consumer (the application is out of memory on its own account), and how the two are told apart from a single signal — here, which container in the pod was killed.
Why the other options fail
- "The product team owns it because the pod is theirs." True of the pod, and it is the answer that leaves the same failure armed in 200 other pods. It also asks a team to debug code they did not write and cannot change.
- "Nobody owns it because the kubelet evicted it." The kubelet enforced a limit somebody chose. Naming the enforcement mechanism as the cause is how a fixable default survives three incidents.
- "Whoever is paged first owns it." This is right during the incident and wrong afterwards. Incident command belongs to whoever is paged; the structural fix belongs to whoever sets the default. Conflating them is how platforms accumulate unowned defects.
Common weak answers
"Raise the memory limit" fixes tonight and nothing else: an unbounded buffer with a bigger bound fails at a higher log rate. The mechanism-level fix is bounded buffering with a documented drop policy, plus a limit derived from measured peak throughput rather than from the injector's first guess. And the answer that scores highest adds the measurement: percentage of pods running each sidecar version, and sidecar-caused restarts per 1000 pods per week, published by the platform rather than discovered by product teams.