advanced 3 min answer

Reddit's 14 March 2023 outage ran 314 minutes after a Kubernetes 1.23 to 1.24 upgrade removed the node-role.kubernetes.io/master label that Calico route reflector selectors depended on, and Kubernetes has no supported downgrade. Sequence the next cluster upgrade so a gate can actually block it.

redditkubernetesupgradeautomated-gatesrollback
Show the full answer Hide the answer

What a gate can and cannot assert

A gate evaluates declarations. The route reflector configuration existed in the running cluster and in the memory of engineers who had left, and in no repository, so no gate and no review board could have asserted anything about it. The first step of this migration is therefore not about the upgrade at all.

The sequence

  1. Make the cluster's own configuration declarative, before touching the version. Export every custom resource, selector, affinity rule and toleration into a repository, then add a CI job that re-exports nightly and fails on any diff against the committed copy. This is reversible, independently useful, and it is the precondition for every step below. Expect it to take longer than the upgrade: a cluster that has run for years holds configuration nobody remembers creating.
  2. Write the gate as a label-contract assertion. For every selector in the exported configuration, check the label keys against the set the target version actually applies. The 1.24 change dropped node-role.kubernetes.io/master in favour of node-role.kubernetes.io/control-plane; a check comparing selectors to the target version's label set finds that mechanically in seconds. No human reading a changelog reliably finds it, which is the point: this class of failure is a machine check, not a judgement.
  3. Prove the reverse path before the forward one. Kubernetes performs schema and data migrations during an upgrade and defines no downgrade, so the rollback is a restore from backup plus a state reload. Reddit's restore procedure had been written against a version that was already end of life and predated the move to CRI-O, so it had to be rewritten during the incident. Make the gate's precondition a dated successful restore of a cluster of this shape within the last 30 days, performed by the on-call rotation rather than by the person who wrote it.
  4. Rehearse on a cluster built from the exported configuration, not on a long-lived staging cluster. Long-lived staging has drifted from production in exactly the undocumented ways that cause this failure.
  5. Order the clusters. You cannot canary 1% of a control plane, so the canary unit is a whole cluster. That is an argument for having more than one, and for upgrading the least critical first with a stated soak period rather than on the same day.

The point of no return

The moment the control plane completes its automatic migrations. Before that, abort. After that, the only path back is the restore, and the time it takes is the outage. Everything in the sequence above exists to move effort to the left of that line.

How long it really takes

The upgrade itself is hours. Steps 1 and 3 are weeks, and they are the migration. A plan that budgets for the upgrade and not for the declaration and restore work is a plan to discover the undocumented dependency in production.

When this is the wrong answer

For a six-node cluster running one stateless service, all of this is heavier than deleting the cluster and rebuilding it from manifests, which is the correct rollback at that size. The apparatus is justified when the cluster is a single point of failure for the product and when its configuration has outlived the people who wrote it.