intermediate 3 min answer

Your CI suite fails about one run in four, and roughly 80% of those failures pass on a re-run. Developers now re-run by reflex, including on real failures. JetBrains ships flaky-test detection in TeamCity based on flip rate. Design the gate, and say what you refuse to do.

jetbrainsteamcityflaky testsquality gatesciquarantine
Show the full answer Hide the answer

The situation

A 25% failure rate where four in five failures are false is not a testing problem yet — it is a signalling problem. The suite still contains the information, and the team has rationally stopped reading it, because the cost of investigating is high and the prior probability of a real failure is 20%. Once re-running is the default response, the gate has been removed without anyone deciding to remove it, and the next real regression will be re-run until it passes or merged past.

What the tooling gives you

TeamCity detects flaky tests automatically by two heuristics, both of which are the right ones:

  • Flip rate — the ratio of status changes (pass→fail or fail→pass) to invocations, measured per agent, per build configuration, over a window that defaults to 7 days. A high flip rate is flakiness almost by definition.
  • Status change on an unchanged revision — a test that passed and then failed with no code change between the builds. This one is near-conclusive: with the input identical, a different output can only come from nondeterminism.

The important property is that both are computed from history the CI system already has. No annotation, no developer judgement, no triage meeting. That matters because the alternative — asking developers to label their own flaky tests — produces labels on the tests they find annoying rather than on the tests that are flaky.

The gate

  1. Classify automatically and continuously. Every test carries a live flip rate. Nothing is labelled by hand.
  2. Split the suite into two gates in one run. Stable tests block the merge. Tests above the flakiness threshold run in the same build, report their results, and do not block. This is the move that restores the signal: a red stable gate now means something, and developers stop re-running by reflex because re-running no longer helps.
  3. A failing quarantined test still creates a ticket, assigned to the owning team, with the flip-rate history attached. Quarantine is a holding pen with a queue, not a bin.
  4. Cap the quarantine, explicitly. "No more than 2% of tests quarantined, and nothing stays longer than 30 days" — then enforce it by failing the build when the cap is breached. Without a cap, quarantine grows until the stable suite tests nothing, which is the standard way this pattern fails.
  5. Fix the cause, not the symptom. Flakiness is overwhelmingly shared mutable state, time dependence, ordering assumptions and real async races. The last category is the valuable one: a test that is flaky because of a genuine race is reporting a production bug, and deleting it discards a finding.
  6. Report the trend, not the count. Flip rate across the suite, weekly. A number that is going down is the only evidence the practice is working.

What I refuse to do

  • Automatic retries on every test. This is the most commonly deployed answer and it is the one that causes the original problem. Retries hide flakiness from the metrics, so the flip rate the detection depends on stops being observable, and the genuine async race is silently masked in CI while remaining live in production.
  • Deleting flaky tests. The flakiest tests are usually the ones touching the most integration surface, which is to say the ones that would catch the most.
  • A manual triage meeting. It does not survive three sprints, and the backlog outruns it from week one.

When this is the wrong answer

Below roughly a few hundred tests, build the discipline rather than the machinery. At that size the team can name every flaky test, and the honest move is to fix them this week — the flip-rate infrastructure, the quarantine lane and the cap are governance for a problem that governance will not solve faster than four hours of work.

The approach also fails if the flakiness is concentrated in the environment rather than the tests: a shared staging database, a rate-limited third-party sandbox, an under-provisioned agent pool. Quarantining tests then just hides an infrastructure problem, and the flip rate will be high across unrelated tests simultaneously, which is the tell. If flakiness correlates with agent or time of day rather than with the test, fix the environment and leave the suite alone.