A mesh runs an Envoy sidecar in each of 2400 pods across 500 services. Every proxy is pushed the whole cluster's endpoints and routes and sits at roughly 50 to 150 MB of memory - on the order of 150 to 350 GB fleet-wide. A proposal replaces sidecars with one shared proxy per node. Which fact about the estate should decide it?
Show the full answer Hide the answer
The deciding property
The memory is config, not traffic. A proxy's footprint in a mesh this size is dominated by the clusters, endpoints and routes it holds, which scale with the size of the mesh rather than with the pod's own request rate. That is why the number is roughly the same on a pod serving 5 requests per second and one serving 5,000.
So the first question is whether the config can be narrowed. Istio's Sidecar resource scopes what an
individual proxy is sent, and mesh-wide discoverySelectors (available since Istio 1.10, documented on
istio.io in 2021) limit which namespaces the control plane watches at all. Typical service graphs are
sparse - a service talks to a handful of others, not to 500 - so scoping removes most of what each proxy
holds. It is a configuration change, reversible in minutes, with no new failure domain.
Why scoping comes first
Replacing sidecars with a per-node proxy is a change of failure domain, and it is not reversible in an afternoon. One proxy process then serves every pod on the node, so a crash or a restart during upgrade affects unrelated workloads that happen to be co-scheduled, and the proxy becomes a shared resource with noisy-neighbour behaviour in connection counts and CPU. That may be an excellent trade, but it should be made against a scoped baseline, not against an unscoped one that overstates the saving several times over.
Two cautions worth stating in the review: scoping config is not a security boundary - istio.io says so explicitly - so it must not be presented as egress control; and the control plane must stay in every proxy's scope or the proxies lose their own configuration source.
Why the other options fail
- p99 latency overhead. A genuine mesh cost and the wrong discriminator here, because a per-node proxy does not remove the hop. It changes where the proxy runs, not how many proxies the request traverses.
- Running two control-plane versions. This is an operational prerequisite for any mesh upgrade, including one you do anyway. It tells you whether you can execute a change, not which change to make.
- Which budget pays. Chargeback changes who complains, not what the fleet consumes. Answering a capacity question with an accounting one is a common and expensive confusion.
- HTTP versus opaque TCP. It constrains which features apply in either topology and does not favour one. It would matter if the proposal were to drop L7 processing entirely, which is not what is on the table.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| Config already scoped and pods per node above roughly 40 | Per-node proxy | Fixed per-proxy overhead now dominates and is paid 40 times per node |
| Workloads are short-lived jobs | Per-node proxy | Per-pod proxy start-up is a large fraction of a job's lifetime |
| Per-workload blast-radius or isolation requirement | Keep sidecars | A shared node proxy puts unrelated tenants behind one process |
| The binding constraint is CPU from TLS and L7 work | Neither | Scoping config does not reduce per-request CPU - that is a capacity problem |
When this is the wrong answer
For a mesh of 12 services in one language, both options are over-built: a competent HTTP client library with retries, timeouts and mutual TLS at the ingress solves the same problems with no new control plane. Scoping config is also the wrong first move if the mesh's real pain is upgrade toil rather than memory, in which case the money goes on making proxy restarts routine.