Alibaba Cloud reported a peak of 583,000 orders per second during Singles' Day on 11 November 2020. Your own platform is far smaller but its annual peak is roughly 40 times a median day and its date is known a year in advance. How does that one fact reshape an eighteen-month migration plan?
Show the full answer Hide the answer
The sequence
A known annual peak is not merely a change freeze. It is a capacity proof deadline and a rollback-invalidation date, and both work backwards from a fixed point rather than forwards from today.
- Place the fence first. Either the cutover completes with at least one full business cycle of soak before the freeze begins, or it moves to after the peak. There is no middle option, because a cutover that lands two weeks before the peak has been proven against ordinary traffic only.
- Decide which system carries the peak, and write it down. New, old, or both. "Both" is a real answer and the expensive one, since it means paying for two capacity envelopes plus the reconciliation between them for a quarter.
- Soak for a cycle, not a fortnight. A month of production traffic exercises the weekly and month-end shapes. Only the peak itself exercises the peak, which is why the fence exists.
- Freeze into a stable coexistence state. Any coexistence machinery that needs a human — a reconciliation job someone inspects, a dual-write divergence report someone clears — must be either automated or stopped before the freeze. During the freeze nobody is available to nurse it.
- Keep the rollback alive across the peak, and cost it at peak volume. Reverse replication that was to be switched off after cutover now has to run through the peak, which is when its lag and its cost are highest.
Where data can diverge, and how you would know
The divergence risk is highest in the days either side of the peak, when volume is abnormal and change is frozen. Run the reconciliation more often during the freeze, not less, and alert on the absolute count of unmatched records rather than a percentage — at 40 times normal volume, a percentage threshold that was sensible in October tolerates 40 times as many broken records in November.
The point of no return
It is not the write switch. It is the moment you release enough of the old system's capacity that it can no longer carry the peak. Decommissioning capacity is usually done for cost reasons, weeks after cutover, by a different team, with no migration decision attached to it. That is the step that should carry a named authoriser and a date, and it should sit after the peak, not before it. Until it happens, rollback is a routing change; after it, rollback is a procurement exercise.
The rollback at each stage
Rehearse the rolled-back configuration under load, not just the forward one. A legacy system that has been serving 5% of traffic for three months has cold caches, shrunken connection pools and capacity reservations somebody has already given away. A rollback path that has not been tested at peak-adjacent volume is a plan to fail twice.
How long it really takes
Working backwards from a November peak with a one-month soak and a one-month contingency, the cutover must complete by the end of August, which means dual-running starts in June and the migration's real deadline is nine months earlier than the calendar year suggests.
When this is the wrong answer
A platform whose load has no calendar should not build its plan around a date. A business-hours enterprise product with smooth growth has no peak worth fencing, and inventing one imposes freeze periods that slow the programme for nothing. There the binding constraint is the release train and the regulator's audit dates. Check whether the peak is genuinely a multiple of the median before letting it govern eighteen months of sequencing.