practice

Zero-Downtime Cutover

Moving a live system to a new implementation or datastore without an outage, using dual writes, backfill and a reversible traffic switch.

migrationavailabilitydata

The sequence is well established and its safety comes from each step being individually reversible.

Replicate: begin copying data from old to new continuously, usually by change data capture, and backfill the history. Verify: compare the two continuously until they agree, and keep comparing. Dual write: the application writes to both, with the old still authoritative — this is the step that requires the most care, since a failure writing to the new store must not fail the transaction, and a failure writing to the old must. Shadow read: read from both, serve the old, compare the results, which surfaces query-level differences before they matter. Switch reads: move read traffic to the new store gradually, with the old still being written and available for instant rollback. Switch authority: the new store becomes authoritative. Decommission the old, after a deliberate soak period.

The properties that make it safe: at every step until the final one, reverting is a configuration change rather than a recovery operation.

The details that break it: writes that are not idempotent, so replay during backfill duplicates; sequence and identifier generation that must not collide between the two; and long-running transactions or batch jobs that hold a view of the old store while the switch happens. Each needs handling explicitly, and each is usually discovered late.