Sidecar Resource Overhead
The cumulative CPU, memory and latency cost of running helper containers alongside every application container.
The sidecar pattern is architecturally elegant — cross-cutting concerns implemented once, in any language, injected everywhere — and its cost is multiplicative in a way that only becomes visible at scale.
A proxy consuming 100 MB and a fraction of a core is negligible per pod and is not negligible across five thousand pods, where it becomes hundreds of gigabytes of memory and hundreds of cores, purchased continuously. Add a logging agent and a security agent and the overhead per pod can rival a small application container.
The latency cost is the other half: traffic traverses a proxy on the way out and another on the way in, adding perhaps a millisecond per hop, which is immaterial for one call and material for a request path with eight hops.
The considerations that follow. Sidecars need their own requests and limits, and under-provisioning them produces application latency that is very hard to attribute. Startup ordering matters — an application container that starts before its proxy is ready fails its first calls, which is a classic source of deployment-time errors. And shutdown ordering matters equally, since a proxy that exits first strands in-flight requests.
The alternative worth weighing for high-density estates is a node-level agent shared across pods, which trades some isolation and per-workload configurability for a large reduction in overhead — the direction several platforms have moved for exactly this reason.