intermediate 3 min answer Multiple choice

A superapp in the mould of Grab renames a column using expand and contract. The rolling deploy of the 240-pod service takes 11 minutes. A batch consumer reads a cached schema descriptor with a 1-hour TTL. A worker fleet is draining a 6-hour backlog of jobs serialised against the old shape. Mobile clients hold the old field name until users update. Roughly how long must both shapes stay correct before the contract step is safe?

grabexpand-and-contractschema-migrationmobile-clientsestimation
Pick one
Show the full answer Hide the answer

The assumptions, stated

The window is a maximum over readers, not a sum. Both shapes must be correct until the slowest live reader of the old shape has stopped reading it. Each reader contributes its own horizon:

  • In-cluster pods: the rolling deploy duration, 11 minutes.
  • The batch consumer: deploy plus one full TTL expiry, so about 71 minutes. A cached descriptor can be up to a TTL old at the moment the deploy finishes.
  • The worker fleet: the backlog drain, 6 hours, plus the retry horizon of any job that fails and is re-enqueued. A job serialised two days ago still carries the old shape when it finally runs.
  • Mobile clients: the adoption curve of the build that stops using the old name.

The arithmetic

max(11 min, 71 min, 6 h + retry horizon, mobile adoption).

Mobile dominates by an order of magnitude. Store-distributed clients commonly reach roughly 80 to 90% of installs within one to two weeks, and the remaining tail runs for months. With a forced-upgrade floor you can bound it: refuse the old build after a stated date and the horizon becomes that date. Without one, there is no horizon you control.

So the defensible answer is roughly two weeks to reach the bulk of clients, and months for the tail unless a minimum-version check cuts it off. Two weeks is the planning number; it is the right order of magnitude and it is the number that changes the plan, because it rules out contracting in the same sprint.

Which assumption dominates the error

Mobile adoption, by far. The server-side terms are minutes and hours and are known within a factor of two. Adoption is weeks and is uncertain by a factor of five depending on whether a forced upgrade exists. Any effort spent refining the deploy-duration estimate is wasted; the only estimate worth sharpening is the client one.

What the number rules in and out

It rules out a calendar-driven contract step. Instead, gate the drop on a measured signal: add a counter on every read of the old column or field, expose it per client version, and drop only after the counter has been zero for twice the longest observed reader interval. That converts a guess into an observation and costs one metric.

It also rules in a cost you must now budget: two shapes correct for two weeks means dual writes, two code paths and two sets of tests for that period, on every service in the chain.

Why the other options fail

  • About 11 minutes. This is the answer if every reader is an in-cluster pod that re-reads the schema on start. It is the most common error because the rolling deploy is the only horizon the deploying team can see on its own dashboard. It ignores every reader that caches, queues or ships separately.
  • About 7 hours. This correctly finds the server-side readers and then sums them instead of taking the maximum, which is a second error hiding inside a right instinct. Summing would matter only if the horizons were sequential; they run concurrently. It also stops at the cluster boundary and forgets that a client is a reader.
  • Indefinitely. True only if you never set a minimum supported version. Treating the window as unbounded is how a schema accumulates permanent dual shapes, and five years later the contract step has still not run. A minimum-version policy is the mechanism that makes the horizon finite.

When this is the wrong answer

If every reader is a stateless in-cluster service with no cache and no queue, the window really is the deploy duration, and planning for two weeks costs a sprint of dual-write work for nothing. Enumerate the readers before sizing the window: the number is an output of that list, not a rule of thumb.