Month 1 of a migration, the reference tables for currency and tax class are copied to the new store and verified row for row. Month 9, a regulator forces a new tax class into the legacy system. Month 11, the new system starts rejecting about 1 order in 400 with a validation error nobody can reproduce in test. What failed, and which design decision made it possible?
Show the full answer Hide the answer
The trigger
The new system validates incoming orders against its own copy of the tax-class list. That copy was correct on the day it was made and has been frozen since. The legacy system gained a tenth tax class in month 9, orders carrying it began flowing in month 11 when the regulation took effect, and the new system rejected every one of them as an unknown code.
The design decision that made it possible: reference data was treated as a one-time migration artefact rather than as a stream with an owner. Transactional data got a replication pipeline and a reconciliation job. The lookup tables got a copy and a tick in a checklist, because they are small and static, and "static" is a property of a snapshot rather than of the data.
Why it propagated
Reference data sits on the validation path of everything. One unknown code does not degrade a feature; it rejects whole transactions, and it does so at a rate set by how common that code is in the business mix — here, roughly 0.25% of orders. The rejections are also correct behaviour by the new system's own rules, so no exception is thrown, no circuit trips, and no error budget is obviously consumed.
The test environment made the symptom unreproducible by construction. It had been seeded from the month-1 snapshot, which is the same snapshot that is wrong in production, so every attempt to reproduce the failure used data guaranteed to be missing the new code.
Why detection lagged
Two thresholds hid it. The rejection returned a generic validation error indistinguishable from genuine bad input, of which there is always a background rate. And 0.25% sat comfortably under a 1% error-rate alert, so the signal existed and was averaged into invisibility for two months.
The structural fix versus the tempting local fix
The tempting fix is to insert the missing row. It takes an hour, it works, and it guarantees the next occurrence, because nothing about the mechanism changed.
The structural fix has three parts:
- Every reference table gets a disposition, in writing, at the start of the migration: either it is continuously replicated from its owning system for the whole coexistence, or ownership transfers on a named date with a change process on both sides. There is no third option, and "copy it and check later" is how the third option is spelled.
- Unknown-code rejections become their own metric, separate from invalid input, alerting on the first occurrence rather than on a rate. A code the system has never seen is a categorically different event from a malformed field, and averaging them together is what cost two months.
- Non-production environments are refreshed from the current reference data, not from the migration-day snapshot, so "cannot reproduce" stops being a property of the test estate.
The general lesson
Anything copied once will diverge over any migration longer than the source system's own change cycle. Do the multiplication before the programme starts: a legacy system whose reference data changes quarterly, against an eighteen-month migration, will produce roughly six divergences. The question is never whether they happen but whether they arrive as an alert or as a customer complaint.
When this is the wrong answer
A migration that completes inside a single change freeze does not need streaming reference data. If the source system is frozen from Friday to Monday and the cutover is inside that window, a one-time copy with row-for-row verification is exactly right, and building replication for lookup tables is machinery with no failure to prevent. The rule is set by the duration of coexistence, not by the size of the tables.