practice

Cutover Abort Criteria

also called Stop Conditions, Numeric No-Go Thresholds

Numeric thresholds agreed before a cutover with one named person empowered to apply them, which replaces an improvised 03:00 judgement made by the people with the most sunk cost in continuing.

migration-riskcutoverrollbackrehearsaldecision-rights

At 03:14 the reconciliation report shows a 0.4% mismatch. Step four of the runbook is at 2.5 times its rehearsed duration. The sponsor is asleep. The runbook says "restore from the pre-cutover snapshot if data integrity is compromised". Nine people have been awake for nineteen hours and the next window is a month away.

"Data integrity is compromised" is not a measurement, so there is nothing to apply. The decision is improvised under fatigue by the group with the most invested in proceeding, and that group reliably chooses to proceed. The distinction from the surrounding artefacts matters: a readiness gate is evaluated before a step and is silent on what the step produces, and a point of no return says where rollback stops being possible. Abort criteria are the numbers that stop a cutover while it is still reversible, which is the only window in which stopping is cheap.

Why it matters

The cost of continuing past a bad signal is not symmetrical with the cost of stopping. Stopping costs a weekend and a rescheduled window. Continuing costs reconciliation of divergent data across two systems, measured in weeks and sometimes impossible to finish.

Fatigue hides the asymmetry. Nineteen hours in, the sunk cost of the night is vivid and the cost of a corrupted dataset is abstract. A threshold written in advance is the only mechanism that survives that, because it converts a judgement into a comparison.

Implementation patterns

  • Express each criterion as a number a dashboard already shows. "Abort if the mismatch rate across 10,000 sampled records exceeds 0.02%" is roughly two records at that sample size and is detectable. "Abort if the new path's error rate exceeds the old path's by more than 0.1 percentage points."
  • Make duration a first-class criterion, such as twice the rehearsed time for any step. Overrun is the earliest honest signal that the plan's model of the system is wrong.
  • Name one decider and one deputy. Not a committee: a committee at 03:00 continues. The authority to abort without consulting the sponsor is granted in writing beforehand.
  • Rehearse at full volume, because a 5% copy measures correctness and not duration, and backfill, index build, replication catch-up and verification are superlinear in volume or dominated by fixed I/O.
  • Pair each criterion with its action: which step to unwind, which flag to flip.

Industry example

The contrast worth studying is the staged way Alipay's transaction database moved onto OceanBase, the distributed relational database that originated inside Alibaba. By OceanBase's own account the plan for the 2014 Singles' Day event was a 1% slice of transaction traffic, and when the incumbent system reached its ceiling that slice absorbed about 10% instead; the rest of the core data links followed over the years after. A ramp of that shape is only a control if each stage has a stated condition for proceeding and for not proceeding: the ramp supplies the reversible position and the criteria supply the decision. The two are close to useless apart, and the public record of failed big-bang banking cutovers is a record of reversibility and stop conditions missing together.

Failure scenarios

  • Qualitative wording, which cannot be applied at 03:00 and functions as permission to continue.
  • The committee. Three people with different incentives and no individual accountability converge on waiting for more information until the point of no return passes.
  • No rehearsed baseline, so "twice the expected duration" compares against an estimate and every overrun is explained away as conservatism.
  • Criteria without a rollback mechanism. The threshold trips, abort is called, and the only rollback is a snapshot restore that discards the first cohort's writes, so the abort is refused.
  • A sample too small to detect the threshold. A 0.02% criterion against 500 records cannot distinguish zero bad records from one, so it never trips.

Trade-offs

Numeric criteria will sometimes abort a cutover that would have succeeded, and that is the price: too tight and the programme never cuts over, too loose and they are decoration.

The calibration that works is to derive each threshold from the cost of what it protects against. A mismatch rate becomes a number of customer records somebody must contact and fix by hand; when that exceeds what operations can absorb in a week, that is the threshold. A criterion derived from remediation capacity is defensible in a review; a round number is not.

When not to use it

A cutover with no irreversible step does not need abort criteria, it needs a rollback button. If the switch is a routing flag and reverse replication runs throughout, the correct behaviour is to flip back on any anomaly and investigate afterwards, and threshold ceremony slows that down.

They are disproportionate for a small change with a bounded blast radius: a single tenant moving between shards, reversible in minutes, is managed by watching and reverting. Abort criteria earn their cost where the step is long, partially irreversible, and performed by tired people when nobody senior is reachable.

Interview question

Q: "Your cutover stages by customer cohort at 2, 10, 50 and 100 percent. Readiness for each stage is 'all tests green and no open P1s'. The rehearsal ran on a 5 percent copy. What is missing, and what would you write instead?"

What a strong answer covers: that a readiness gate is not an abort criterion, because it is evaluated before the step; three or four thresholds stated as numbers with their sample sizes, including duration against a rehearsed baseline; one named decider with written authority; why a 5% rehearsal measures correctness and not duration; the pairing of each criterion with an executable rollback, since criteria without a mechanism get overruled; and what to leave alone, since the cohort ramp is the best control in the plan.

Quick check

Quiz: Why is "abort if data integrity is compromised" worse than no criterion at all? — Because it looks like a control and cannot be applied, so the decision falls to fatigued people with sunk cost in continuing while the document reassures reviewers that a stop condition exists.

Flashcard: What three properties does a usable abort criterion have? — A number a dashboard already shows with a sample size large enough to detect it, a baseline measured in a full-volume rehearsal, and one named person with written authority to apply it alone.