Flip Rate
also called Test Flip Ratio, Status Change Rate
The ratio of a test's pass-to-fail and fail-to-pass transitions to its total invocations, which identifies flakiness from history the CI system already holds rather than from anyone's judgement about which tests are annoying.
A suite fails one run in four, and four in five of those failures pass on a re-run. The team has correctly calculated that the prior probability of a real failure is 20% and that investigating costs an hour. So re-running becomes the reflex, and the gate is gone without anyone deciding to remove it.
Fixing it requires knowing which tests are flaky, and the instinct — ask developers to label them — produces labels on the tests people find irritating rather than the nondeterministic ones.
Flip rate solves the identification problem mechanically. A flip is a status transition for one test between consecutive runs: pass to fail, or fail to pass. The flip rate is flips divided by invocations, over a window. A deterministic test that is correct flips zero times. A deterministic test that is broken flips once, when it breaks. Only a nondeterministic test flips repeatedly, so a high ratio is close to a definition of flakiness rather than a proxy for it.
Why it matters
The metric needs no new instrumentation and no human input. Every CI system already stores per-test results per build. Flip rate is a query over data that exists, which is why it can be applied to a 30,000-test suite on the day someone decides to.
It also separates two things teams conflate. A test with a 40% flip rate and one that has failed on every run since Tuesday look the same on a dashboard of red tests and need opposite responses: nondeterminism versus a regression. Flip rate distinguishes them automatically.
Its second-order value is as a trend: aggregate flip rate across the suite, weekly, is the only credible evidence a flakiness effort is working. Quarantine counts measure activity; the trend measures outcome.
Implementation patterns
- Compute it per agent and per build configuration, not globally. TeamCity does this deliberately, for a diagnostic reason: a test flipping only on one agent is an environment problem, and the global number hides that.
- Window it. Seven days, TeamCity's default: long enough for a signal, short enough that a fixed test leaves the list.
- Pair it with the stronger signal: a status change on an unchanged revision. A test that passed and then failed with no code change between builds is near-conclusive, because identical input with different output can only be nondeterminism. Use this as the high-confidence classifier and flip rate as the ranking.
- Split the suite into two gates in one run. Stable tests block the merge; tests above the threshold report without blocking. This is what restores the signal — a red stable gate means something again, and re-running stops helping, so people stop doing it.
- Cap the non-blocking lane and enforce the cap. "No more than 2% of tests, nothing longer than 30 days", with a build failure when breached.
- Ticket every quarantined failure to the owning team, with the flip-rate history attached. Quarantine is a holding pen with a queue.
Industry example
JetBrains documents flaky-test detection in TeamCity on exactly these two heuristics: a flip rate — status changes relative to invocation count, measured per agent, per build configuration, over a 7-day default window — and a status flip on a build with no changes. The design decision worth copying is that the product computes this and surfaces the tests, rather than offering a way for users to tag them. It removes the step that does not survive contact with a busy team, which is the step where a human is asked to maintain a list.
Failure scenarios
- Blanket automatic retries. The most commonly deployed response to flakiness, and it destroys this metric specifically: a retried test reports as passing, so the flips stop being recorded and the flip rate goes to zero while the nondeterminism remains. The flakiness is hidden from the measurement rather than removed, and a genuine async race stays masked in CI and live in production.
- Quarantine without a cap, which grows until the blocking suite tests nothing and the gate is theatre.
- Treating environment flakiness as test flakiness. High flip rates across unrelated tests at once, or correlated with agent or time of day, mean a shared database, a rate-limited sandbox or an under-provisioned agent pool. Quarantining tests then hides an infrastructure problem.
- Deleting the flakiest tests. They touch the most integration surface, so they would catch the most — and one that flakes from a real race is reporting a production bug.
- Flip rate on a rarely-run test. Invoked twice a week, there is no meaningful denominator and the ratio is noise.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Flip-rate quarantine lane | The blocking gate becomes trustworthy again within days | Quarantined coverage is genuinely lost until each test is fixed |
| Automatic retries | Green builds immediately, no triage work | Flakiness becomes unmeasurable and real races are masked |
| Fix on discovery, no lane | No coverage lost, root causes addressed | Only viable while the team can hold the whole flaky set in its head |
When not to use it
Below roughly a few hundred tests, build the discipline instead of the machinery. At that size the team can name every flaky test, and the honest move is to spend an afternoon fixing them. A flip-rate pipeline, a quarantine lane and a cap are governance for a problem governance will not solve faster than the direct fix.
It is also the wrong instrument when flakiness is environmental rather than test-level, and the tell is in the metric itself: high flip rates spread across unrelated tests, or clustered on one agent. Reading that as a list of tests to quarantine treats the symptom and leaves a shared dependency or an overloaded runner in place.
Finally, never use flip rate as a target: driving it down with deletions or retries improves the metric and worsens the suite. Manage the count of root causes fixed.
Interview question
Q: Your CI fails 25% of the time, 80% of those failures are false, and developers now re-run by reflex. You have one sprint. What do you do?
What a strong answer covers: recognising this as a signalling problem first, and that the gate is already effectively gone · classifying automatically from existing history rather than asking people to label · splitting into blocking and non-blocking lanes in one run as the move that restores the signal · refusing blanket retries, because they make flakiness unmeasurable and mask real races · capping the quarantine in the build · and checking whether flakiness correlates with agent or time of day, which would make the whole plan wrong.
Quick check
Quiz: Why does a blanket retry policy make flip rate useless? — Retried tests report as passing, so status transitions are no longer recorded; the ratio falls to zero while the nondeterminism is unchanged.
Flashcard: How many times does a deterministic but broken test flip? — Once, when it breaks. Repeated flipping is close to a definition of nondeterminism, which is why the ratio identifies flakiness rather than merely correlating with it.