Legitimate Divergence
also called Expected Parallel-Run Difference, Correct-on-Both-Sides Delta
The class of differences between an old and a new analytical platform that are correct on both sides, which a parallel run must classify and budget for rather than investigate as defects.
Forty reports are run on both the legacy warehouse and its replacement. Thirty-one match. Nine differ, by amounts between 0.02% and 0.4%. The migration stalls for six weeks while engineers chase each one, and at the end roughly a third of them turn out to be cases where both platforms computed exactly what they were asked to.
Legitimate divergence is the set of differences produced by correct code on both sides, arising from platform-level semantics that were never part of anyone's specification: numeric types, timezone interpretation, collation, null ordering and rounding mode. It is not a rounding inconvenience to be waved away. It is a category that has to be named, expected and budgeted before the parallel run starts, because the alternative is an investigation with no finishing condition.
Why it matters
A parallel run is the main risk control in a warehouse migration, and its value depends entirely on being able to close it. An unclassified difference is indistinguishable from a defect, so without a taxonomy the team investigates everything at the same priority, the schedule slips, and the political cost of the migration rises until someone proposes running both platforms permanently.
There is a second-order effect that decides careers. Some of the differences are defects in the legacy platform that have been producing a slightly wrong number for years. Discovering that is the most valuable output of a migration and the most awkward, and a team that has not agreed in advance how such findings are handled will suppress them.
Implementation patterns
- Write the tolerance per report before the first comparison. A revenue total may need to match to the penny; an engagement rate may match to 0.5%. Both are legitimate positions and neither can be decided while a difference is on screen.
- Classify before investigating. Route each difference to one of five buckets: numeric type, timezone, collation or sort, rounding, or unexplained. Only the last bucket gets an engineer.
- Pin types in the model rather than arguing about engines. Cast money to a fixed decimal explicitly and division results too, so the two platforms cannot disagree about precision.
- Normalise timezone at the boundary. Store timestamps with an explicit zone and convert once, so a day boundary is a property of the data rather than of the session running the query.
- Compare distributions, not only totals. Two platforms whose totals match can still disagree about which rows are in which bucket, and a totals-only comparison passes it.
- Keep the classification log. It becomes the evidence pack for sign-off and for the auditor who asks why a restated figure changed.
Industry example
The pattern is generic across platform families rather than specific to one vendor, which is why it surprises teams repeatedly. A legacy appliance-era warehouse commissioned around 2008 that stores currency as a fixed decimal, and a cloud engine of the 2020s that promotes a division to double precision, will disagree in the last places of a per-row calculation; summed across tens of millions of rows the disagreement becomes visible at the fourth significant figure. Neither engine is wrong. The specification never said which one it wanted, because in a single-platform world there was nothing to specify.
Failure scenarios
- The endless parallel run. No tolerance was agreed, so every difference is open, and the cutover date moves until the migration is cancelled or the old platform becomes permanent.
- Tolerance set after seeing the number. The team picks a threshold that makes the current difference acceptable, which destroys the control's credibility with finance and audit.
- A real defect hidden in the noise. A genuine join error producing a 0.3% difference is closed as "type promotion" because that was the previous nine explanations.
- Legacy errors buried. A discovered defect in the old platform is quietly matched in the new one so the reports agree, and the organisation carries the wrong number forward permanently.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Tight tolerance everywhere | Maximum confidence; finance signs readily | Long parallel run; engineering spent on differences nobody would notice |
| Tolerance by report criticality | Finishes; effort lands where the money is | Requires a judgement per report and someone accountable for it |
| Match the legacy platform bit for bit | No arguments at cutover | Ports legacy defects into the new platform as requirements |
When not to use it
When the migration is a re-implementation rather than a move, the concept does not apply and insisting on it will wreck the project. If the business has agreed that definitions change, that the new platform computes revenue differently because the old computation was wrong, then a report-by-report comparison is measuring two different things and every difference is expected. Replace the parallel run with acceptance tests against newly specified expected values. The same applies to reports being retired: comparing a report nobody will use after cutover is effort spent on a deliverable with no future.
Interview question
Q: You are three weeks into a parallel run. Nine of forty reports differ by less than half a percent and the business sponsor is asking whether the new platform can be trusted. How do you answer, and what do you do next?
What a strong answer covers: refusing a yes or no and offering a classification instead; naming the specific semantic causes rather than "rounding"; proposing a tolerance per report agreed with the report owner and recorded; separating the differences that are defects in the new build from those that are defects in the old one, and naming who decides what happens to the second group; and setting a date by which unclassified differences either become defects or become accepted, so the run has an end.
Quick check
Quiz: Name three causes of a parallel-run difference where both platforms are computing correctly. - Numeric type promotion such as decimal against double precision; timezone interpretation moving rows across a day boundary; collation, case sensitivity or null ordering changing which rows match or which row a first-value function returns.
Flashcard: Why must the tolerance for each report be written down before the first comparison is run? - Because a tolerance chosen after seeing the difference is indistinguishable from an excuse, and the parallel run loses its value as evidence with finance and audit.