Data-Plane Upgrade Skew
also called Proxy Version Skew Window, Mesh Upgrade Skew
The window during which a mesh runs proxies of more than one version alongside a control plane that supports only a bounded range of them, which is what makes every mesh upgrade a scheduled fleet restart rather than a deployment.
A platform team schedules a mesh upgrade for a Tuesday afternoon and finds, three weeks later, that 340 of 2400 pods are still on the old proxy because nothing has restarted them, the control plane cannot be upgraded until they move, and a security fix is waiting behind both.
A mesh has two software fleets with different deployment mechanics. The control plane is a deployment: you roll it and it is done. The data plane is one proxy per pod, and a proxy version changes only when the pod restarts, which happens when the workload owner deploys, or when something evicts it. The platform team does not control the schedule of its own upgrade — that is the property everything else in mesh operations follows from.
The skew window is the interval in which both versions are live. It is bounded above by how much version difference the control plane supports with its proxies, and bounded below by how quickly the fleet's pods naturally recycle. When the first bound is smaller than the second, the platform must force restarts, and that is a conversation with every product team rather than a change.
Why it matters
The size of the window sets the cost of every mesh security patch and every mesh feature. With a supported skew of one version, a two-version jump is two fleet restarts of every pod, in order, not one. A 2400-pod fleet restarted in tiers with a day between tiers and a week of soak for connection-lifetime changes is six to ten weeks of elapsed time and a few engineer-weeks of attention, repeated every time the mesh releases something you need. This recurring bill, not the per-proxy CPU and memory, is what makes a mesh expensive for a small estate.
The window is also where the subtle regressions live. Two proxy versions can differ in retry semantics or header handling in ways that are correct in isolation and inconsistent between callers, so a symptom appears on some paths and not others, in proportion to how far the restart has progressed.
Implementation patterns
- Publish proxy version per pod per namespace as a metric before the first upgrade. A mesh upgrade you cannot see the progress of is indistinguishable from one that stalled.
- Pin the injected version per namespace so the version a pod gets is a platform decision, and rollback is a restart in the other direction rather than a redeploy of the control plane.
- Move connection lifetime first. Envoy's documentation is explicit that on hot restart existing connections are not transferred: they complete during the drain or are terminated. In Kubernetes the sidecar upgrade is a pod restart, so the lever is a maximum connection duration rolled out weeks earlier, which teaches long-lived gRPC clients to reconnect on a schedule they already tolerate.
- Set drain timers for internal traffic. Envoy defaults to a 600-second drain and a 900-second parent shutdown with a gradual strategy, which suits an edge proxy; the documentation itself suggests values such as 60 and 90 seconds for service-to-service use.
- Restart in tiers: platform test services, internal low-tier, read paths, then write paths and stream holders. One tier per day at most, because the failures you are hunting are slow.
- Watch from outside the mesh. Success rate and p99 from the caller's client library, not from the proxy telemetry, because the emitter is what changed.
Industry example
Envoy was created at Lyft and open-sourced in 2016 to move networking concerns out of application code, and its hot-restart machinery exists precisely because replacing a proxy under live traffic is the recurring operational cost of the model. The documented behaviour is the honest teaching point: the new process initialises fully, takes copies of the listen sockets from the old process over a unix socket, and the old process drains — and connections that do not finish inside the drain are cut. Every mesh built on that data plane inherits the property, which is why "upgrade the mesh" and "restart the fleet" are the same sentence.
Failure scenarios
- A stalled long tail. Services that deploy rarely never restart, so they hold the control plane hostage; the security patch waits on the least active team.
- A reconnect storm. Upgrading a tier that holds hours-long streams without having rolled out a connection-duration limit resets them together, and the burst looks like an application fault.
- Skew exceeded. The control plane is upgraded first for convenience, older proxies stop accepting configuration, and endpoint changes silently stop propagating — traffic keeps flowing to endpoints that no longer exist.
- Retry amplification during the window. A default retry change between versions doubles offered load on a dependency for the duration of the rollout, visible as retry volume rather than as errors.
Trade-offs
Shrinking the window by forcing restarts buys a predictable upgrade and spends product teams' goodwill and their error budgets. Letting the fleet recycle naturally costs nothing and leaves you unable to state when the fleet will be patched. Most mature platforms take a middle path: a published maximum pod age, enforced by eviction, so restarts are routine and nobody negotiates during an upgrade.
When not to use it
This whole apparatus is the argument against a mesh for a small, single-language estate. A dozen services sharing one HTTP client library get retries, timeouts, mutual TLS termination at the edge and metrics with no per-pod proxy and no fleet restart per release. Adopt the mesh, and this operational model with it, when the estate is polyglot and large enough that per-language libraries drift faster than a proxy fleet can be restarted — in practice tens of services across three or more runtimes.
Interview question
Q: You must take 2400 sidecars from version n-2 to n, the control plane supports one version of skew, and several services hold gRPC streams open for hours. Walk me through the sequence and tell me where the point of no return is.
What a strong answer covers: version inventory as a metric first; two fleet restarts rather than one because of the skew bound; connection-duration limits rolled out weeks earlier so streams recycle; drain timers chosen from p99 duration rather than the edge defaults; tiered restarts with per-namespace version pinning so rollback is a restart; the point of no return being the control-plane move to n after which the old data plane is unsupported; and measurement from the caller's side because mesh telemetry is the thing under test.
Quick check
Quiz: Why is a mesh upgrade a fleet restart rather than a deployment? — Because a proxy version changes only when its pod restarts, so the platform's upgrade schedule is set by every workload owner's deploy cadence unless restarts are forced.
Flashcard: What bounds the skew window in a mesh upgrade and what happens when it is exceeded? — The control plane's supported version range; beyond it older proxies stop accepting configuration and endpoint changes stop propagating silently.