metric

Exposure Lag

also called First-Execution Delay, Dormancy Window

The interval between code reaching production and that code first executing on real input - the window in which a release looks proven by live traffic although the change it carries has never run.

fastlylatent-defectbake-timecanaryfeature-flags

A deployment finishes, error rate is flat, the bake passes and the release is signed off. Four weeks later a routine customer configuration change takes most of a network down, and nothing was deployed that day.

Exposure lag is the gap between the two events people treat as one: the artifact arriving in production, and the first execution of the path the change introduced with inputs that reach it. While the gap is open, every signal a rollout gate reads describes code that has not run. Bake time measured from the deploy proves the deploy; it cannot prove a branch no request entered.

Why it matters

Rollout safety rests on the assumption that production traffic exercises what you shipped. For the ordinary change it holds. It fails for exactly the changes that hurt: a feature behind a flag nobody has enabled, a new construct in a configuration language, an error-handling path, a recovery routine.

The arithmetic makes the hole visible. A canary at 1% for 30 minutes on a service doing 2,000 rps covers roughly 3.6 million requests, and zero of them execute a branch gated behind a switch no customer has turned on. The gate's statistics are computed over a population that excludes the change.

Detection is unaffected by the lag. Attribution is destroyed by it, because the change record that caused the incident is weeks old and the timeline starts with today's innocent-looking change.

Implementation patterns

  • First-execution counters per new path or flag, emitted once per path per version, so the pipeline records "shipped at T, first executed at T plus x".
  • Gate on exercise rather than elapsed time: a stage advances when the new paths have accumulated executions, and otherwise holds instead of counting minutes.
  • Synthetic activation in the canary — a smoke suite that enables each newly shipped flag once and sends the inputs that reach it, which collapses the lag from weeks to minutes for paths you can name.
  • A flag inventory with a maximum age. A branch never executed 90 days after shipping is either dead code or an armed landmine.

Industry example

Fastly, 8 June 2021. A software deployment that began on 12 May carried a bug triggerable by a specific customer configuration. On 8 June a customer pushed a valid configuration change containing that condition, and about 85% of the network began returning errors from 09:47 UTC. Fastly's account records detection within about a minute, the configuration disabled by 10:27, and 95% of the network healthy within 49 minutes.

The exposure lag was about 27 days: four weeks in which a deployment looked proven by the largest traffic sample available anywhere, while the defective path waited for an input nobody had sent.

Failure scenarios

  • Misattribution. The incident is recorded against the configuration change, the latent defect survives the postmortem, and the class recurs with a different customer.
  • The false bake. Canary analysis reports no regression because the new code contributed no samples, which reads identically to a clean result.

Trade-offs

Choose Gains Pays
Gate on first-execution counts the bake measures the change rather than the clock per-path instrumentation and rollouts that stall when real traffic never arrives
Synthetic activation in canary lag falls to minutes for named paths a suite that must track every new flag and traffic that pollutes production metrics
Accept the lag and invest in fast reversal no new machinery nothing to deploy against a silent or delayed failure

The decision rule: when the failure class is loud, fast reversal is the cheaper control; when it is silent or delayed, close the lag instead. One-minute detection is why reversal worked for Fastly, and a wrong rounding rule found by tomorrow's reconciliation is the case where it would not.

When not to use it

Do not measure it for every change. A stateless transform with no flags, where every request exercises every line, has a lag near zero by construction, so instrumenting it teaches nothing. Skip it also when you cannot name the new paths: a counter nobody maintains decays into a stale inventory, which is worse than none because it suggests the paths are known.

Interview question

Q: Your canary ran for 30 minutes at 5% of traffic and the release was clean. Three weeks later a routine configuration change causes a major outage, traced to that release. What should have been different, and what will you measure from now on?

What a strong answer covers: that the canary measured the deployment and not the change; that the metric is time from ship to first execution, instrumented per path; that the gate should advance on executions rather than minutes, with synthetic activation where traffic will not arrive; and the limit, which is that closing the lag pays only where failures are silent.

Quick check

Quiz: A release has been live 20 days with no regression and its new path's flag has never been enabled. What evidence do you have about that path? Answer: none; the flag's first enable is the real first day of the rollout.

Flashcard: What does bake time measured from the deploy fail to prove? — Anything about a code path production traffic never entered; the proving event is first execution, not deployment.