practice

Test Quarantine

also called Flaky Test Isolation

Removing an unreliable test from the blocking suite immediately, so it stops eroding trust, while keeping it running non-blocking until it is fixed or deleted.

flakinesstrustcisignalpolicy

A flaky test — one that passes and fails on the same code — is not a weak test. It is a negative-value test: it provides no assurance, and it removes the assurance of every reliable test around it.

The mechanism is behavioural. Once a suite fails intermittently for unrelated reasons, the rational response to any failure is to re-run. Everyone learns this quickly, and from that point a genuine failure is also re-run, re-run again, and eventually merged — because it looks like the flakiness everyone has learned to ignore.

Quarantine restores the signal immediately by removing the flaky test from the blocking set, without losing the information it might still provide.

The policy

  1. Detect automatically from pass/fail history. A test that fails and then passes on the same commit is flaky by definition — this is measurable rather than anecdotal, and the measurement is what makes the policy enforceable.
  2. Quarantine on detection, immediately and without debate. It continues to run, non-blocking, gathering data.
  3. Fix or delete within a defined window — typically a week or two. Quarantine is a holding state, not a destination; a quarantine population that only grows is the same problem with an extra step.
  4. Attribute to an owner, because an unowned flaky test is never fixed.
  5. Track the flakiness rate as a headline metric, so the trend is visible and the worst offenders are known.

The causes worth checking first

Most flakiness is architectural rather than a test defect, which is why fixing the injection points removes whole categories at once:

  • Shared mutable test data, making tests order-dependent and defeating parallelism.
  • Timing assumptions — fixed sleeps, races between an action and its observation.
  • Real external dependencies, failing for reasons outside the team's control.
  • Non-deterministic inputs — randomness, wall-clock time, ambient configuration.
  • Resource contention under parallel execution.

Industry example

Suites dominated by end-to-end tests exhibit this most severely, because any component, any timing issue and any environment problem fails the test. The correct treatment there is stronger: a flaky end-to-end test should be deleted rather than quarantined, since its value was already low and its cost is high.

The broader lesson is that flakiness is a leading indicator of testability problems. A team with persistent flakiness usually has code that cannot be tested deterministically, and the durable fix is injecting time, randomness and environment rather than stabilising tests one at a time.

Failure scenarios

  • Tolerating flakiness, which trains everyone to re-run and destroys the suite's signal.
  • Quarantine with no expiry, becoming a permanent graveyard.
  • Retry-on-failure as the policy, which hides flakiness and also hides genuine intermittent defects — some of which are real production races.
  • Blaming test authors when the cause is non-injected dependencies.
  • No flakiness measurement, so the problem is argued about rather than tracked.

Trade-offs

Quarantining removes coverage, and occasionally the quarantined test would have caught a real defect. That is a genuine cost and it is smaller than the alternative — a suite nobody trusts provides no coverage at all, regardless of what it nominally tests.

Automatic retry is tempting because it is cheap, and it is the wrong trade: it converts a visible problem into an invisible one and masks real intermittent defects.

Interview question

"A test in your suite fails about one run in twenty and the team has learned to re-run it. What is that test actually costing you, and what would you do this week?"