A team plans to move 2.5 TB to a new database by taking the application offline at 22:00 on Saturday, restoring a backup into the new store, and switching over by 08:00 Sunday. The data copies fine in a test at one tenth the volume. Why does this plan usually fail before it is executed?
Show the full answer Hide the answer
The arithmetic that kills it
The window is ten hours. Fill it with what actually has to happen.
- Restore. A 2.5 TB restore sustaining roughly 150 to 250 MB/s takes about 3 to 5 hours, and sustained throughput is what matters rather than the peak the test showed on 250 GB.
- Index and statistics rebuild. On a loaded relational store this is commonly comparable to the load itself, so allow another 1 to 2 hours before the database answers queries at a usable speed.
- Verification. Row counts plus checksums on the largest tables, and a smoke test of the application against the new store: 1 to 2 hours if it is scripted and much longer if it is not.
That is 5 to 9 hours of a 10-hour window before anyone decides anything. The rollback, which is restarting the old system and re-running its own smoke test, needs perhaps an hour and has to start before the window closes. The plan has no slack and therefore no decision point, which means the team will discover a problem at 06:30 with the only remaining option being to press on.
The rule
A cutover window must contain the work, the verification, and the rollback, in that order, with the rollback's start time fixed in advance. Write it as a clock time, not as a condition: "if the application smoke test has not passed by 05:00, we restore service on the old system." That single line converts a hope into a plan, and it is the line most often missing.
Why the other options fail
- Backup format incompatibility. A real constraint when moving between different engines, and it is a reason to use a logical export or a replication tool rather than a physical restore. It does not apply when both sides run the same engine, and it changes the method rather than the feasibility.
- Writes lost after the snapshot. This is the classic failure of snapshot-based migration, and the stem has already excluded it: the application is offline for the whole window, so no writes are arriving. Reach for this answer when the plan says "take a backup on Saturday and cut over next week", where it is fatal.
- Restore cannot be parallelised. Modern restores do parallelise across files and tables, and where they do not the answer is more throughput rather than a different plan. It changes the hours in the estimate, not the absence of a rollback budget.
When not to build the migration machinery
Below roughly 200 GB, with a rehearsed restore and a business that genuinely tolerates a Sunday outage, the big-bang window is the cheapest correct answer. Dual writes, backfill, shadow reads and a reversible traffic switch are the right machinery for a system that cannot stop, and they cost weeks of engineering that protect nothing when the restore finishes in 20 minutes and the rollback is "start the old service".
The deciding number is not the data size on its own. Choose the window unless the rehearsed work plus verification plus rollback exceeds about half the window, at which point the extra complexity of an online migration is buying you the only thing the window cannot give: the ability to stop. Public incident write-ups since 2018 on failed weekend cutovers repeat the same shape — the team was in production with no reversal budget left and pressed on because the alternative was an unplanned Monday outage.