concept

Leading Health Signal

also called In-Band Invariant Signal, Detection-Latency Gate

A rollout gate input whose latency is shorter than the damage it indicates - the property that decides whether any bake time can catch a defect, as distinct from how much traffic the rollout exposes.

progressive-deliveryinvariantsbake-timeobservability-latencyrollout-gates

A pricing change rolls out in nine steps over 90 minutes with a 10-minute bake between them. Error rate, latency and saturation stay flat. The defect is a wrong rounding rule that under-charges some orders, found the next morning when finance reconciliation runs.

Every rollout control worked. Nothing was wrong with the bake time, the step sizes or the gates. The problem is that the defect's first evidence did not exist until roughly 14 hours after the damage began, and a gate can only act on signals that arrive inside its bake window.

That quantity deserves a name. Observability latency is how long after damage occurs the first evidence exists, and a gate is meaningful only if its signal's latency is shorter than the window you can afford. A signal that satisfies that is leading; everything else is trailing.

Why it matters

Progressive delivery is usually discussed in terms of exposure: what fraction of traffic, which cohort, how long between steps. Those control how much damage a defect does and say nothing about whether you will find out, which signal latency settles.

With a 90-minute rollout and 14-hour detection, the rollout completes about 13 hours before the first evidence, and at 1,200 orders a minute that is on the order of a million orders priced wrongly. Extending the bake to cover it means nine days per release, which adds more risk through batch size than the longer bake removes. You cannot buy safety with time here; the only move is a signal that arrives sooner.

Implementation patterns

  • An in-band invariant on the write path. Assert a property the system must satisfy regardless of the code under test: for a charge, total equals the sum of line items plus tax to the cent.
  • Independence is the whole requirement. The check must compute its expectation with a separate implementation. Calling the same function and comparing proves only that it agrees with itself.
  • Gate on violations per thousand writes, so the threshold holds across steps of different sizes.
  • Shadow comparison where an invariant cannot be stated, at roughly double the per-request work.
  • Shorten the trailing signal where you cannot lead. Reconciliation every 15 minutes instead of nightly turns a 14-hour latency into a 15-minute one.

Industry example

Automated canary analysis as practised since the mid-2010s made rollout decisions statistical rather than human, comparing a canary against a concurrent control on error rate and latency. It is the mature form of the exposure half of the problem and is limited to the signals it has.

The complement is the invariant checking long used in payments and ledger systems, where a transaction is validated against an independently computed expectation before commit and a mismatch counter is a first-class alert in production. Put the two together and a rollout gate can read a signal that exists in milliseconds.

Failure scenarios

  • A bake time set by habit, 10 minutes because the previous team used 10 minutes, with no statement of the fault it should surface.
  • A noisy invariant disabled within 14 days, leaving the gate weaker than before.
  • Cellular rollout treated as detection: blast radius is divided and the latency is unchanged, so one cell is damaged for 14 hours instead of all cells for 13.

Trade-offs

An invariant costs CPU on every request, a few percent for arithmetic checks and about double the work for a full shadow execution, and it is code that can itself be wrong. You can only check properties you can state, so a price that is arithmetically correct and wrong by policy passes.

Against that, it is the only mechanism that turns an unobservable rollout into a gated one. Decision rule: build leading signals for money, identity and deletion, and accept trailing detection everywhere else.

When not to use it

If the damage is cheap to reverse, a trailing check plus a forward fix is the better trade. A wrong value in a reporting field, backfilled in an hour, does not justify a permanent per-request cost.

Do not reach for an invariant when the real problem is exposure rather than detection. If the signal already arrives in seconds and a minority cohort is hidden inside an average, the fix is cohort-level analysis against a concurrent control, not a new write-path check.

Interview question

Q: "A nine-step rollout with bake times showed nothing, and finance found the bug the next morning. Tell me what was wrong with the rollout design, and what you would change without slowing releases down."

What a strong answer covers: naming observability latency and comparing it to the affordable bake window; the arithmetic showing the rollout finished hours before any evidence; why extending the bake trades one risk for a larger batch-size risk; the in-band invariant with independence as the binding requirement; the CPU and false-positive costs; limiting the practice to money, identity and deletion; and the distinction from cellular rollout, which divides damage without detecting it.

Quick check

Quiz: A rollout's gates look fine and the defect is found 14 hours later. Which quantity was never checked? — The observability latency of the gate's signals, against the affordable bake window.

Flashcard: What property decides whether a write-path invariant proves anything? — Independence. If it computes its expectation with the code it is checking, it only proves the code agrees with itself.