advanced 3 min answer

You must move the sidecar proxies of 500 services and 2400 pods from mesh version n-2 to n. The control plane supports one version of skew with its data plane and each proxy upgrade is a pod restart. Several services hold gRPC streams open for hours. Plan the migration so that no service sees an unplanned reset and you can stop at any point.

service meshenvoyupgradeversion skewdraining
Show the full answer Hide the answer

The sequence

  1. Publish the version inventory first. Proxy version per pod per namespace, as a metric, before anything moves. A mesh upgrade is a fleet restart; if you cannot see the version spread you cannot tell a stalled upgrade from a finished one.
  2. Upgrade the control plane to n while the data plane is still n-2 only if the documented skew allows it. With one version of supported skew it does not, so the real sequence is n-2 to n-1 across the data plane, then the control plane to n, then the data plane to n. Two fleet restarts, not one. Teams who plan for one discover this halfway through.
  3. Change the connection lifetime settings before the upgrade, not during it. Long-lived streams are the hard part: a proxy restart cannot hand over established connections. Envoy's documentation is explicit that on hot restart existing connections are not transferred — they must complete during the drain or be terminated. In Kubernetes the sidecar upgrade is a pod restart, so the equivalent lever is a maximum connection duration that forces clients to reconnect on a schedule they already tolerate. Roll that out first and watch reconnect rates for a week.
  4. Set drain timers for service-to-service traffic. Envoy's defaults are a 600-second drain and a 900-second parent shutdown with a gradual drain strategy, which suit an edge proxy. The documentation itself suggests much shorter values such as 60 and 90 seconds for service-to-service deployments; pick from your own p99 request duration plus the longest stream you have agreed to break.
  5. Restart in tiers, pinning the injected version per namespace: platform test services, then internal low-tier, then read paths, then write paths and anything holding streams. One tier per day at most, because the failures you are looking for are slow.

Where it can diverge and how you would know

The mesh's own telemetry is the wrong place to look for mesh regressions, because the emitter is what changed. Watch from outside it: request success rate and p99 per service from the caller's client library, connection resets per second, and configuration acceptance per proxy version. A silent divergence to expect is a policy or retry default whose semantics changed between versions, which shows up as a small change in retry volume rather than as an error.

The rollback at each stage

Because the injected version is pinned per namespace, rollback is the same restart in the opposite direction, and it is available up to the point where the control plane moves to n and the old data plane version is no longer supported. That is the point of no return, and it is why the control plane moves after the data plane has been proven at n-1.

How long it really takes

Two fleet restarts of 2400 pods, tiered, with a day between tiers and a week of soak for the connection-lifetime change: six to ten weeks of elapsed time and a few engineer-weeks of attention. Mesh upgrades are quarterly work that never finishes, which is the part of the bill teams do not price when they adopt one. Envoy itself came out of that pressure: Lyft built it and open-sourced it in 2016 to move networking concerns out of applications, and its hot-restart machinery exists because upgrading proxies under live traffic is the recurring cost of the model.

When this is the wrong answer

With a dozen services in one language, the mesh is the wrong tool and this whole plan is the argument against it: a shared HTTP client library with retries, timeouts and metrics has no fleet restart. Tiered proxy upgrades are justified when the fleet is polyglot and large enough that per-language libraries drift faster than a proxy fleet can be restarted.