advanced 3 min answer

Your CI system holds cluster-admin credentials for 14 production clusters and deploys by applying manifests at the end of each pipeline. You must move to in-cluster reconcilers pulling from Git, with no change freeze and no loss of the ability to deploy while the migration runs. Sequence it.

gitopsreconciliationcredentialsfield-ownershipcutover
Show the full answer Hide the answer

The sequence

  1. Make the repository describe what is actually running. Export live manifests per cluster, commit them, and run a server-side dry-run diff on a schedule until the diff is empty for the fields the repository claims. Comparing whole objects never converges: defaulted fields, webhook-injected sidecars, autoscaler-owned replica counts and operator-written spec fields live only in the cluster. Adopting server-side apply and deciding field ownership is step one, not a detail.
  2. Install the reconciler in diff-only mode on every cluster, alerting on drift, with sync disabled. Both systems are live, CI still deploys, and you now have a measurement of how wrong the repository is.
  3. Enable sync for one low-traffic namespace in one cluster. Its pipeline step changes from an apply to a commit plus a wait-for-reconciled check, and CI's apply for that namespace is removed in the same change, so no namespace has two writers.
  4. Expand namespace by namespace, platform add-ons last, because a reconciler managing the networking add-on it depends on can remove its own prerequisites.
  5. Replace CI's cluster-admin credential with a read-only one once the last namespace is synced, keeping the admin credential in the secret manager's version history.
  6. Build and test a break-glass path outside both CI and the reconciler, time-limited and audited, because every deploy now depends on the reconciler being healthy.

Where data can diverge and how you would know

The hazard is step 3's overlap: if CI's apply and the reconciler both touch a namespace, the last writer wins and the flapping is silent. Make the apply step refuse when it sees the reconciler's ownership marker, so the overlap is a loud failure rather than a quiet fight.

Two signals carry the migration: per-resource drift from the reconciler, and an audit-log query for writes by any identity that is not the reconciler. The second is the honest measure of how much of your change surface the pipeline covers, and the number is usually a surprise.

The point of no return

Not the credential removal — that reverses in minutes from the secret manager. The real one is when the only way to change production is a merge, which arrives the moment you delete CI's apply step and have not yet tested break-glass.

The rollback at each stage

Steps 1 and 2 are free, being read-only. Steps 3 and 4 roll back per namespace by suspending the reconciler for it and restoring the CI apply step, which is a revert of one commit. Step 5 rolls back by reinstating the credential.

How long it really takes

The per-cluster work is days; the per-service ownership cleanup is the schedule. With 60 services, expect roughly 1 to 2 days each of deciding who owns which field, so the calendar cost is 60 to 120 engineer-days of other teams' attention. For 14 clusters and a two-person platform team, a quarter is realistic and the gating item is ownership cleanup, not tooling. A plan that quotes 2 weeks has assumed step 1 away.

What you buy: no outside system holds admin credentials, manual changes revert within one reconcile interval of a few minutes, and the deploy path survives a CI outage. What you pay: the reconciler is now a dependency of every change, so its own upgrades are production changes.

When not to migrate at all

One cluster and a small team, or a release that is genuinely a workflow: run a migration job, wait, deploy, warm a cache, flip a flag. A reconciler is a convergence loop, not an orchestrator, and expressing ordering inside it means writing an operator to recover sequencing a pipeline gave you for free. Keep push where you cannot run a privileged agent at all, such as deploying into a customer's cluster, and where one cluster has run in production for years with a single pipeline identity that is already auditable.