advanced 3 min answer

At 14:05 a canary on 5% of traffic passes automated analysis across 60 metrics and is promoted. At 14:40 the fleet error rate is 12 times baseline. The canary's own dashboards were clean the entire time. What made the canary blind, and which design decision allowed it?

canaryprogressive deliverystatisticssilent failurerollback
Show the full answer Hide the answer

The trigger

A regression that fires on a narrow slice of requests, promoted by a gate that could not have detected it. The interesting failure is not the code change. It is that the release process produced a confident pass on evidence that was never capable of saying no.

Why the canary could not see it

Three mechanisms, each independently sufficient.

The comparison was against the wrong population. Canary analysis compared the new instances to the running fleet, and the fleet has been up for days: warm caches, warm JIT, full connection pools, settled garbage collection. A freshly started instance differs from that fleet for reasons that have nothing to do with the change. To stop those differences firing constantly, someone widened the thresholds, and a noise band wide enough to absorb cold-start variance is wide enough to absorb a real regression. The fix is a baseline control: deploy the old version at the same moment onto the same number of equally fresh instances, and compare canary to control. The cost is one extra deployment group per release, and it is the cheapest thing in this list.

Sixty metrics is sixty chances to be wrong. At a 5% significance threshold, roughly three metrics signal falsely on every clean run. A gate that cries wolf three times per release gets overridden, so it either stopped blocking or stopped being read. Prefer four or five gating metrics — error rate, latency, saturation, one business metric — and keep the other 55 as diagnostics with no authority. A gate that can block must be a gate someone trusts in production at 02:00.

The event arithmetic never worked. 5% of 4,000 requests a second is 200 a second. A fault firing once in 50,000 requests produces an event every four minutes in the canary, so 35 minutes yields about nine events, against a baseline error rate of one in a thousand. Nine events cannot be distinguished from noise at that base rate. Canary duration is set by how many events of the kind you are looking for will accumulate, not by minutes on a clock. If the signal cannot reach significance at 5%, either raise the share or state plainly that the canary does not cover that class of fault.

Why detection lagged after promotion

Metrics were aggregated across versions. Without a build identifier as a dimension on error and latency series, a regression in the new version appears as a fleet-wide rise with no attribution, and the first 20 minutes go on deciding whether it is the deploy at all.

The structural fix versus the tempting local fix

The tempting fix is "run the canary longer", which addresses one of the three mechanisms and leaves the wrong-population comparison and the multiple-comparison problem exactly as they were.

The structural fix is four changes: a same-age control group, a short list of gating metrics, a version dimension on every series, and automatic rollback keyed to the canary-versus-control comparison rather than to a human reading a dashboard. The last one matters most, because every minute of human interpretation is a minute of full exposure after promotion.

When not to trust a canary at all

A canary is an experiment, and its power is traffic share times duration times base rate. An experiment without a control group is an anecdote, and a gate with a 5% false-positive rate on 60 metrics is a gate that will be disabled by the people it inconveniences. State the fault classes your canary cannot detect, in writing, next to the ones it can.