advanced 3 min answer

An Envoy-based mesh carries 600 services across 3000 pods. At 09:00 the mesh control plane is lost entirely and cannot be restored for 90 minutes. No application or node has failed. What happens over those 90 minutes, what sets the real deadline, and what would make the mesh ride it out?

service-meshenvoyxdsfail-staticcertificates
Show the full answer Hide the answer

Minute by minute

09:00 to 09:01. Steady-state traffic is unaffected. Envoy latches onto the last configuration it received from the management server and retries the connection in the background, so listeners, clusters and routes keep working from the proxy's in-memory copy. The mesh is fail-static by design, and the documented xDS behaviour is what makes the next 90 minutes survivable at all.

09:01 onward, anything new is broken. A pod created after 09:00 has a sidecar with no configuration. It cannot route outbound traffic, and depending on injection settings it either starts serving without policy or never becomes ready. Three ordinary mechanisms create new pods: deploys, horizontal autoscaling, and node replacement. At 3,000 pods with even 2% hourly churn that is about 60 dead pods an hour, concentrated in whichever service was mid-rollout.

09:05 onward, endpoints go stale. Terminating pods stay in the proxies' endpoint lists until a new push arrives, and no push is coming. Callers send requests to addresses that no longer answer and depend on retries and outlier ejection to work around it. Error rates lift a little; p99 lifts more, because each failed attempt costs a connect timeout.

Any config change is now a silent no-op. Traffic shifts, timeout changes and new routes merge, show green in the pipeline and take no effect. This is the most dangerous part of the window, because the organisation believes it has made a change.

What sets the real deadline

Not the 90 minutes. It is the earlier of two clocks.

  • Certificate lifetime. Workload certificates are short-lived by design; Istio's default is on the order of 24 hours with rotation at about half-life. A proxy that cannot reach the certificate authority stops being able to renew, and once its certificate expires, mutual TLS fails closed and the service is down. That gives something like 12 hours of headroom, not 90 minutes.
  • Pod churn. If anything forces mass pod replacement during the window (a node pool upgrade, an autoscaler event, a crash loop) the fail-static property gives you nothing, because fail-static protects pods that already have configuration.

If you have configured an xDS resource TTL so proxies drop resources after a period without contact, that TTL replaces the certificate clock and can be far shorter. It is a correctness feature with an availability price.

What stops it, concretely

Freeze the things that create pods: disable autoscaler scale-up, pause deploys, hold node-pool operations. Then make the dangerous mode visible rather than silent: configure the injected sidecar to hold the application container until the proxy has configuration, so a pod without policy fails to become ready instead of serving traffic unprotected. Finally, alert on the proxies' connected-state statistic, not on the control plane's own liveness probe, because the question that matters is whether the data plane is still receiving pushes.

What would have to be true for it to self-heal

Certificate lifetime longer than the outage, no pod churn in the window, no xDS TTL, and a control plane that can be rebuilt from declarative state with no mesh dependency of its own. A control plane that needs the mesh to come up cannot come up, which is the circular dependency to design out before you need it.

When this is the wrong thing to worry about

Below roughly 50 services in one language and one cluster, the mesh is carrying cross-cutting concerns that a shared HTTP client library would carry with no control plane to lose. The right answer there is not a resilient mesh; it is no mesh.