pattern

Ambient Data Plane

also called Node-Level Mesh Proxy, Shared Node Data Plane

Running mesh functions in one shared proxy per node instead of a proxy in every pod - so cost scales with nodes rather than workloads - at the price of making the node the failure domain.

service-meshsidecaristioztunnelblast-radius

A platform with 2,000 pods adds a sidecar proxy to each one. At 50 to 100 MB of memory per proxy that is 100 to 200 GB of RAM committed before a single request is served, and the arithmetic is worse than the total suggests: a 30 MB proxy beside a 40 MB service nearly doubles the pod, so the overhead lands hardest on the small services that a mesh was supposed to make cheap to run.

Sidecars also couple lifecycles. The proxy must be ready before the application's first outbound call, drained after its last, and told to exit when a batch job finishes — three ordering problems that every mesh adopter meets in week one.

The ambient alternative moves the data plane out of the pod and onto the node: one shared proxy per node handles connection-level concerns for every pod on it, and layer-7 features are added selectively through a separate proxy per service or namespace rather than universally.

Why it matters

The cost model changes shape. Mesh overhead becomes a function of node count, not workload count, so the hundredth small service on a node is nearly free. Istio's ambient mode, which reached general availability in version 1.24 in 2024, implements this with a per-node proxy called ztunnel for layer-4 identity, mutual TLS, authorisation and telemetry, plus optional waypoint proxies where layer-7 policy is actually wanted; the project reports resource savings that can exceed 90% in favourable cases.

What that buys is adoption. Teams that refused a mesh because of per-pod cost or injection mechanics can get identity and encryption without touching their pod specification at all.

Implementation patterns

  • Split the layers deliberately. Layer 4 — identity, mutual TLS, authorisation, telemetry — in the node proxy; layer 7 — retries, header routing, traffic splitting — only where a waypoint exists.
  • Keep identity per workload, not per node. The node proxy acts on behalf of each pod, which means it holds credentials for every workload on the node and sits in the trust path for all of them.
  • Capture traffic at the node, with redirection rules outside the pod network namespace, so no application change and no injection is required.
  • Plan node proxy capacity explicitly: its CPU is shared, so a chatty workload competes with its neighbours for the proxy's cycles and for its connection table.
  • Audit which services actually need layer 7. If the answer is most of them, the waypoint fleet plus the node proxies can cost more than sidecars did.

Industry example

Istio's ambient mode is the reference implementation, with ztunnel written in Rust specifically to keep the per-node footprint small, and the project's own announcement framing the change as removing the need to over-provision memory and CPU per pod. Cilium takes a comparable approach with a per-node proxy. Both are public projects whose design documents state the trade-off openly, which makes them a better source than any vendor benchmark.

Failure scenarios

  • Blast radius moves from one pod to the node. A node proxy crash or bad configuration push affects every pod on that node, typically tens of them, where a sidecar fault affected one.
  • A layer-7 policy that is silently not applied, because the service has no waypoint. The configuration exists, the retry does not happen, and nothing errors.
  • Noisy-neighbour contention on proxy CPU, which appears as latency in an unrelated service on the same node.
  • Upgrades become node-level events, so a proxy upgrade is now a drain-and-roll exercise rather than a rolling restart of one workload.
  • Compliance gaps where an auditor expects per-workload traffic isolation and the shared proxy is a shared process.

Trade-offs

Choose Gains Pays
Node-level proxy Cost scales with nodes; no injection; no lifecycle coupling Node-sized blast radius; shared CPU; layer 7 only where added
Per-pod sidecar Per-workload isolation and full layer 7 everywhere Memory and CPU per pod; startup and shutdown ordering

When not to use it

When nearly every service needs layer-7 policy, because then you pay for both tiers and the saving disappears. When per-workload process isolation is a stated control rather than a preference. And when the fleet is small: at 50 pods the per-pod overhead you are optimising is a rounding error, and the right answer is a library or no mesh at all. The threshold worth remembering is that sidecar cost becomes a budget line somewhere in the high hundreds to low thousands of pods.

Interview question

Q: Your platform runs 3,000 pods with sidecars and the memory bill is now visible to finance. You are offered a node-level data plane. What do you ask before agreeing, and what will you have to tell your security reviewer?

What a strong answer covers: how many services genuinely use layer-7 policy today, because those need waypoints · the change in blast radius from pod to node, and the node drain procedure that follows · that workload identity is still per-pod but the node proxy holds all of it, so node compromise is a broader event · the migration path, namespace by namespace, with both modes coexisting · and the measurement that settles it, which is memory and CPU per node before and after rather than a vendor's percentage.

Quick check

Quiz: What does a node-level mesh proxy change about failure domains? A proxy fault stops being per-workload and becomes per-node, affecting every pod scheduled there.

Flashcard: Mesh cost scales with pods under sidecars. What does it scale with in ambient mode, and what is the catch? — With nodes, so small services become nearly free; the catch is a node-sized blast radius, shared proxy CPU, and layer-7 features applying only where a waypoint has been added.