Service Mesh Platform  ·  View 21 of 31  ·  5 · Runtime

Total Control-Plane Outage

istiod, SPIRE server and OpenBao are all down in one cluster. Traffic freezes on last-known-good for 12 hours, pages before 24, and fails closed after.

Editable source SVG draw.io All views
0 – 12 h 12 – 24 h Beyond 24 h When it returns Existing traffic Flows on last config Flows · expiry paging Fails closed per pod Renew · reconnect Certificate renewal Not yet due Renewals fail, retried Emergency issuance Backlog issued ≤ 20 s Routes and denies Frozen at last push Frozen Frozen Push · drift clears New pods Refused in strict ns Refused in strict ns Refused in strict ns Config · SVID · ready Crashed proxy Pod unready · drained Pod unready · drained Pod unready · drained Restart fetches config Emergency deny Cilium L3 policy only Cilium L3 policy only Cilium L3 policy only Mesh deny lands Total Control-Plane Outage — A Freeze, Then a Deadline istiod, SPIRE server and OpenBao all down together. 12 hours is the promise; 24 hours is the hard edge set by certificate lifetime. v 1.0 · owner Platform Networking Architecture · date 2026-09

Decisions

  • Proxies keep their last acknowledged configuration indefinitely. There is no staleness timer, because the certificate already bounds how long a proxy can run alone and a second timer would only fail first.
  • A crashed proxy does not come back until istiod does. Envoy keeps configuration in memory, and persisting xDS to disk would mean forking the proxy agent. The pod goes unready and its peers carry its share.
  • An emergency deny during the outage is a Cilium network policy. It is coarse and it does not depend on istiod.

The promise, stated plainly

  • Existing traffic: no degradation for 12 h. Between 12 and 24 h, renewals fail and pages fire while traffic still flows. Beyond 24 h, pods whose certificates expire fail closed one by one.

Deviation from the requirement

  • The requirement asks that a crashed proxy restart from a local config cache. This design does not provide that during a total outage, and says so: the cost is the pods that crash in that window, and the alternative is maintaining a fork of the mesh's agent indefinitely.