pattern

Per-Entity Migration State

also called Record-Level Ownership Flag, Entity Migration Flag, Per-Row Routing State

A field on each business entity naming the system currently authoritative for it, which makes the unit of migration one row rather than one capability - so traffic moves in nameable cohorts and rolls back by updating a flag.

coexistence-patternscohort-migrationroutingdual-ownershiprollback

A capability-level switch has no smaller unit than the capability. Routing "order creation" to the new system moves every order for every customer at once, which is why teams with only capability-level routing end up with a cutover rather than a ramp, and why their first stage is 100%.

Per-entity migration state puts the routing decision on the row. Each customer, account, tenant or order carries a field naming the system that owns it now, and every read and write path consults that field before choosing a backend. Moving 2% of customers becomes an update of a few thousand rows, and rolling them back is the same update in reverse. That is what makes a cohort ramp possible: each stage's blast radius is a population you can name, contact individually and remediate by hand. Without per-entity state, "2%" can only mean 2% of requests chosen at random, which leaves every customer partly on each system.

Why it matters

Coexistence periods run for a year or more, and what makes them survivable is a reversible position at every point. A flag is the cheapest reversible position available: a data change, not a deployment, not a restore, not a configuration push to a fleet. It also gives the programme its only honest progress metric, because "70% of customers are authoritative on the new system" is a count from a table while "the migration is 70% complete" is an estimate, and a stalled strangler migration is usually one where nobody could produce the first number.

Implementation patterns

  • Flip the flag and copy the entity's data in one transaction, on the system losing ownership. This is the load-bearing rule. Flip first and the new system serves an entity whose data has not arrived; copy first and both systems accept writes for the same entity.
  • Read the flag inside the read, in the same query or transaction, rather than caching it per request. A flag cached for a long request routes the write to a backend that no longer owns the row.
  • Keep one authoritative location for it: a column on the entity's own row in the legacy store early on, and a routing service once the new system owns most entities.
  • Make it an enum with an in-flight value, not a boolean. legacy, migrating, new lets a writer reject rather than guess during the copy, which is a visible failure instead of silent divergence.
  • Reconcile per flag value. Entities marked new should have no recent legacy-side writes, an assertion that is cheap, continuous and catches the worst failure directly.

Industry example

The closest documented systems are sharding migrations that deliberately separate the routing level from the physical layout. Figma's horizontal sharding work, described in March 2024, split logical sharding from the physical failover and rolled the logical layer out through database views first, so the irreversible step was short; the first sharded table shipped in September 2023 with about 10 seconds of partial write availability on the primaries. The shared idea is a routing decision expressed as data and changed independently of where the data physically lives. Notion's scheme shows the limit of the analogy: its hash into one of 480 logical shards is deliberately immutable, because that stability is what makes resharding configuration.

Failure scenarios

  • Flag and data flip separately. The entity is marked new, the copy fails, and the new system serves an empty record while both sides diverge with no error anywhere.
  • Both sides writable for one entity, the classic result of copy-then-flip, usually found weeks later once both versions hold real business changes.
  • A cross-entity transaction. A refund crediting a customer owned by the new system and cancelling an order line owned by the old one has no commit boundary. This is the hard limit, not a bug.
  • The flag becomes permanent, a schema dependency of both systems and every report.
  • A consumer that does not read the flag, serving the wrong half of the population silently.

Trade-offs

Choose Gains Pays
Per-entity state Cohort ramps; rollback by data update; a countable progress metric A lookup on every read path; a schema dependency in both systems; no commit boundary across differently owned entities
Capability-level routing One place to change; no per-row invariant No unit smaller than the capability, so every stage is large
Dual write everywhere No routing decision at all Two writes with no transaction between them and reconciliation as the only correctness mechanism

When not to use it

Do not use it when a business transaction spans entities whose ownership can differ. Per-entity routing then creates cross-system transactions you cannot commit, and the right move is to draw the ownership boundary around whatever must commit together, even if that forces a larger first stage. Do not use it for a population small enough to move in one step either: ten thousand accounts migrable in a maintenance window do not need a year of per-row routing.

And do not use it when you cannot change every read path. A single consumer that bypasses the flag makes the pattern worse than no routing at all, because the system looks controlled while part of the population is served wrongly.

Interview question

Q: "You need to migrate a customer base in cohorts of 2, 10, 50 and 100 percent over an eighteen-month coexistence period, and roll any cohort back within minutes. Design the routing, and tell me the one invariant that makes it safe."

What a strong answer covers: a per-entity ownership field as the routing input, with each cohort defined as a query over it; the invariant, that the flag flip and the data copy commit in one transaction on the system losing ownership; reading the flag inside the read rather than caching it; an in-flight state so writers reject rather than guess; reconciliation asserting no recent legacy writes for migrated entities; and the hard limit of a transaction spanning entities with different owners. A strong answer names the exit: the field is dropped, and a plan that never drops it has not finished.

Quick check

Quiz: What must commit in the same transaction as copying an entity to the new store, and why? — The entity's ownership flag, on the system losing ownership. Separating them means either the new system serves data that has not arrived, or both systems accept writes for one entity and diverge silently.

Flashcard: Why does a cohort ramp need per-entity state rather than percentage-based request routing? — Because a percentage of requests leaves every customer partly on each system, so no cohort can be named, contacted or rolled back and reconciliation means nothing.