Your CI retries a failed test up to three times before reporting a stage red. Trunk has been green for eleven straight weeks. Twice this quarter a bug reached production that an existing test had in fact failed on, and p50 pipeline duration is unchanged. Where do you look, and what is the diagnosis?
Show the full answer Hide the answer
The first three things I would look at, and why in that order
- The retry records for the last 90 days: how many stages passed only on attempt two or three, and which tests. This is the whole investigation if the data exists.
- The per-test distribution of attempts. A test that needs a retry on 1 run in 20 behaves differently from one that needs it on 1 in 4, and the second is almost never infrastructure.
- Whether retried failures cluster on a code path or on a resource. Failures that follow a feature are product defects; failures that follow a registry pull or a DNS lookup are environmental.
The diagnosis
Auto-retry cannot distinguish a flaky test from an intermittently failing product. A race condition in the application that fails one run in four has a 25% chance of failing any single attempt, so with three retries the stage passes with probability 1 − 0.25³, about 98%.
The retry converted a bug that reproduces one time in four into a pipeline that reports green 98% of the time. The test did its job twice this quarter and the pipeline overruled it.
The misleading signal
Eleven weeks of green trunk reads as health, and it is the anomaly. On a codebase with this much concurrent change, a gate that never rejects anything is not a gate with nothing to reject; it is a gate that is not discriminating. Flat p50 duration reinforces the illusion, because retries only lengthen the runs that would otherwise have been red, which barely move a median.
The fix
- Retry is permitted but recorded. A stage that passed only on retry reports a distinct state, not green, and the test that needed it is counted.
- Quarantine on a threshold. A test exceeding roughly 1% retry-only passes over 200 runs leaves the gate for a quarantine suite with a named owner and a deadline. Quarantine is a holding pen with an expiry, not a graveyard.
- Never retry an assertion that touches concurrency or shared state without first trying to reproduce it deterministically. Those are the ones most likely to be real.
- Never retry the stage that gates the deploy. The further right the retry sits, the more it is suppressing the signal you are about to act on.
The cost is honest: counting retries will turn a green board amber for a while, and someone has to own the backlog that appears. That is the price of a gate that means something.
The alert that would have caught it earlier
Retry-only pass rate, per test, per week, with a threshold that pages the owning team rather than the platform team. One metric, and it ranks the work for you.
A second, cheaper one: the count of tests that have never failed in 12 months. Combined with the first, it tells you which parts of the suite are noise and which are inert.
When not to remove the retry
For genuinely external flakiness, retry is correct and the alternative is worse. A registry returning 503, a transient DNS failure, a cloud API rate limit: retry those, and retry the infrastructure step rather than the assertion, so the blast radius of the retry is the fetch and not the verification. The distinction is mechanical: retry things outside your code, never the check on your code.