Backfill Governor
also called Throttled Backfill, Health-Gated Migration
A control loop that sets a migration's write throughput from live database health signals and keeps a resumable cursor - so a backfill measured in days cannot turn into an outage and pausing it costs nothing.
Expand-and-contract adds a column in one deployment and removes the old one in another. Between them sits the step nobody sizes: populating hundreds of millions of existing rows. A single UPDATE holds locks, bloats the table and generates write-ahead log faster than replicas can apply it.
A backfill governor makes that work a background process with a throttle. The migration runs in bounded batches, records a cursor after each one, and yields whenever the database says it is unhealthy — so "how fast can we backfill" is answered by the database rather than by a human. The reframing matters more than the mechanism: a backfill is a capacity decision with a duty cycle, not a step inside a deployment.
Why it matters
The damage from an unthrottled backfill does not appear where people watch. A Postgres update is copy-on-write: each row writes a new heap tuple plus entries in every index whose pages change, so a 2 kB row can produce several kilobytes of WAL. The primary absorbs it and the replicas apply it serially, so the first user-visible symptom is stale reads from read replicas, which nobody attributes to a migration.
The sizing follows. At 1,000-row batches of about 100 ms, full speed is 10,000 rows/s, so 420 million rows takes roughly 12 hours; at a 50% duty cycle it is a day, and at 2,000 rows/s about 2.4 days. Both schema shapes must stay correct throughout, so the contract step belongs in a later release.
Implementation patterns
- Batches and sub-batches. The batch bounds the transaction; the sub-batch bounds the statement, so no query exceeds the statement timeout.
- A cursor record per migration, so a pause, deploy or crash resumes rather than restarts.
- Idempotent writes, usually a conditional update or upsert, so a retried batch is harmless and no row is counted twice.
- Health signals that drive the throttle: write-ahead log backlog, replica apply lag, autovacuum on the target table, and a database apdex. Trip any one and pause for minutes rather than slowing marginally.
- Adaptive batch sizing from recent jobs' durations, with a concurrency cap, a kill switch and published progress, so the plan has a date rather than a hope.
Industry example
GitLab documents a batched background migration framework for its own schema changes. Migrations run in batches with a sub-batch size, and the framework pauses when database health indicators trip — the documented signals include pending write-ahead log, autovacuum running against the affected table, and the Patroni apdex falling below its service level objective. Batch size is adjusted from recent jobs' efficiency, a job that times out can be split, and a small number of migrations run in parallel.
Take the shape rather than the thresholds: the throttle is driven by the database's own signals, not a fixed sleep chosen by the migration's author.
Failure scenarios
- Replica lag as a product outage. The backfill runs flat out, replicas fall minutes behind, and users see their own writes vanish on the next page load.
- The non-resumable job. A deploy restarts the worker at 70% and the migration begins again from row one, which is how two days becomes two weeks.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Aggressive throughput | the dual-shape window closes sooner | replica lag, vacuum debt and a real risk of a user-visible incident |
| Health-gated throttle | the migration cannot cause an outage | an unpredictable finish date, so planning treats it as a background task |
When not to use it
When the value can be derived on read. A fallback expression in the query plus population on next write removes the backfill entirely, at the cost of a slightly more complex read path until old rows age out. Choose the backfill only when the column must be indexed or aggregated, when the derivation is expensive, or when a regulator needs the stored value.
Skip the machinery for small tables: a few hundred thousand rows inside the statement timeout is one transaction. And if your platform's online schema-change tool already throttles on measured replica lag, use it rather than writing a second control loop with different thresholds.
Interview question
Q: You must populate a new column on a 420 million-row table that is written to continuously, and the column is required before you can drop the old one. Walk me through the plan, the signal you throttle on, and how you would know it is safe to run the contract step.
What a strong answer covers: batches with a resumable cursor and idempotent writes; throttling on write-amplification signals rather than primary CPU; an estimate of one to three days with the duty cycle named as the dominant uncertainty; the contract precondition of every reader being on the new shape, including cached schema descriptors and draining queues; and the cheaper alternative of deriving the value on read.
Quick check
Quiz: Which signal should pause a backfill, and which looks fine throughout? Answer: pause on write-ahead log backlog and replica apply lag; primary CPU stays comfortable while replicas are minutes behind.
Flashcard: Why is a backfill not a deployment step? — Its duration is set by a duty cycle the database controls, measured in days, so it needs a cursor, idempotent writes and a progress metric rather than a place in a release pipeline.