A food-delivery platform in the mould of Zomato rolls out a pricing change in nine steps over 90 minutes with a 10-minute bake between steps. Error rate, latency and saturation stay flat the whole time. The defect is a wrong rounding rule that under-charges some orders, and it surfaces the next morning when finance reconciliation runs. Which change makes this class of rollout safe?
Show the full answer Hide the answer
The deciding property
A rollout gate can only act on signals that arrive inside the bake window. The property that settles this question is the defect's observability latency: how long after the damage occurs the first evidence exists. Here it is roughly 14 hours, because the only observer is a nightly reconciliation job.
The rollout finishes about 13 hours before the first evidence. At 1,200 orders a minute, that is on the order of a million orders priced wrongly before any gate could fire. No affordable bake time closes a 13-hour gap, so you cannot buy safety with time. You have to shorten the signal.
Why an in-band invariant works
An invariant check runs on the write path and asserts a property the system must satisfy independently of the code under test. For a charge, that is arithmetic: assert that the total equals the sum of line items plus tax to the cent, computed by a separate small implementation, and emit a counter when it does not.
Latency falls from 14 hours to milliseconds, and the counter is now a rollout gate input. A gate on "invariant violations per thousand writes" would have failed at step one, at roughly 1% of traffic.
The cost is specific and bounded. A recomputation of an arithmetic invariant typically costs a few percent of request CPU. The hard requirement is independence: if the check calls the same pricing function, it proves only that the function agrees with itself.
What it costs and what it misses
Invariants catch violations of properties you can state. They do not catch "the price is legal but wrong by policy", so this is not a replacement for review. Each invariant is also code that can be wrong, and a false-positive invariant that blocks rollouts gets disabled within two weeks.
Decision rule: build invariants for money, identity and deletion, and accept trailing detection for everything else. Those three classes are where the damage is irreversible or regulatory, which is what justifies the engineering.
Why the other options fail
- Extend the bake to 24 hours. It would work, and it is unaffordable. Nine steps at 24 hours is nine days per release, which collapses deployment frequency and makes every change a large batch, raising risk by more than the longer bake removes. Long bakes are correct for rare high-stakes changes, not for a pricing engine that ships weekly.
- Compare the canary against a concurrent control. This is the right technique for a different problem: it removes environmental noise so a small regression in a signal you have becomes detectable. Here the problem is that the signal does not exist yet in any form, so comparing two populations on error rate and latency compares two sets of flat lines.
- One cell at a time. Cellular rollout limits blast radius, which is genuinely valuable, and it does nothing about detection. With a 14-hour observability latency you would damage one cell for 14 hours rather than all of them for 13, and in money terms you have divided the loss, not prevented it.
When this is the wrong answer
If the damage is cheap to reverse, a trailing check plus a forward fix is the better trade. A wrong value in a reporting field, caught by tomorrow's job and backfilled in an hour, does not justify a write-path invariant and the permanent CPU cost it carries. Reserve this for writes you cannot take back.