Expand and Contract Migration
also called Parallel Change, Expand-Migrate-Contract
Changing a schema through a sequence of individually backward-compatible steps, so old and new application versions can run simultaneously.
A rolling deployment means old and new application code run at the same time against one database. Any schema change that only one of them tolerates is an outage. Expand and contract removes the problem by never making such a change:
- Expand. Add the new structure in a backward-compatible way — a nullable column, a new table, an additional index. Old code is unaffected.
- Migrate. Deploy code that writes both old and new. Backfill historical rows in batches. Both representations are now correct.
- Transition. Deploy code that reads the new representation. Verify.
- Contract. Stop writing the old, then remove it — often a separate release, days or weeks later.
Every step is independently deployable and independently reversible, which is the entire point.
Why it matters
It converts the hardest coordination problem in a shared-database organisation into a protocol. Without it, schema changes require sequencing releases across teams, which drops the whole organisation's velocity to that of its slowest participant — and makes people avoid necessary migrations, so debt accumulates because change feels dangerous.
Implementation patterns
- Never rename; add and remove. A rename is simultaneously a create and a drop, so it cannot be backward compatible. Add the new column, dual-write, migrate reads, drop the old.
- Never add
NOT NULLwithout a default in one step. Add nullable, backfill, then add the constraint. - Batch backfills with rate limiting, so a migration does not become a load incident on the primary.
- Dual-read with comparison during transition — read both, use the old, log discrepancies. This turns the migration's correctness into a measurement rather than a hope.
- Feature-flag the read switch, so cutting over and rolling back are configuration changes rather than deployments.
- CI enforcement. Lint migrations for blocking DDL on large tables and for destructive changes without a deprecation period. Reviews catch what people remember; pipelines catch what they do not.
Industry example
The pattern is essential wherever several teams write to the same core tables and change schemas independently — the situation in most large enterprise estates, and in fast-growing platforms whose database predates their team structure.
The characteristic failure it prevents is precise: a team adds a NOT NULL column with a default; another
team's nightly batch insert, written before the column existed and specifying its columns explicitly,
fails at 2 a.m. Nobody involved did anything unreasonable, and the changing team could not have listed the
affected consumers.
Expand and contract makes that impossible by construction, because at no point does the schema require the new column to be present. Combined with a consumer registry derived from query logs, it removes most of the fear that makes organisations slow at schema change — and fear, rather than technical difficulty, is usually the binding constraint.
Failure scenarios
- Contracting too early, before every deployed version has stopped using the old structure — including long-running batch jobs, offline workers and rolled-back releases.
- The permanent expand. Dual-write is left running for years because nobody schedules the contract, so the schema accumulates half-finished migrations and the dual-write code becomes load-bearing.
- Backfill without rate limiting, saturating the primary.
- Forgetting non-application writers — reporting tools, ETL jobs, direct support corrections.
- Assuming the ORM handles it. Most migration tools will happily generate a single-step destructive change.
Trade-offs
The cost is real: three to four deployments instead of one, temporary dual-write code, a period where two representations must be kept consistent, and the discipline to finish. For a small system with a single deployable and an acceptable maintenance window, a stop-the-world migration is simpler and entirely legitimate.
The pattern earns its cost when downtime is unacceptable, when multiple teams deploy independently, or when the table is large enough that a blocking change is measured in hours.
Interview question
"You need to change a column's type on a table with 500 million rows, in a system that deploys twelve times a day with no maintenance window. Walk me through every deployment, and tell me what you would monitor between each one."