Flaky Test Management
Quarantine, detection, and the trust a suite loses once red stops meaning broken.
7 to work through
-
beginner Multiple choice
One integration test fails only in CI, only on Tuesdays, only in the 00:30 UTC scheduled run. It passes locally and on every pull-request run. What do you look at first?
2 min answer -
intermediate
A CI pipeline runs 2,400 tests in 18 minutes and fails about one run in four for reasons unrelated to the change. To stop the interruptions, the team enables automatic retry: any failing test is re-run twice, and the build is green if it passes on any attempt. What happens over the following six months?
3 min answer -
intermediate
A suite has persistent flaky tests. Why is tolerating them worse than deleting them, and what policy works?
2 min answer -
intermediate
A team's CI suite fails roughly one run in three for reasons unrelated to the change. Everyone reruns until green. How do you recover the situation?
2 min answer -
intermediate
A test suite has a 4% flake rate. The team re-runs failed builds. What is the actual cost, and what policy fixes it?
2 min answer -
intermediate
Write the policy for handling flaky tests. What are the rules?
2 min answer -
intermediate
Your test quarantine has grown to 140 tests over a year. What has gone wrong and how do you recover?
2 min answer
4 terms in this topic
Flake Budget
A hard limit on test flakiness, enforced by automatic detection and quarantine with a deletion deadline - because above a small rate the suite stops …
conceptFlake Rate
The proportion of test failures that are not caused by a real defect - a quantity with a sharp threshold, above which the suite stops being a signal …
conceptFlaky Test
A test that passes and fails on unchanged code, destroying the signal value of the entire suite it belongs to.
practiceTest Quarantine
Removing an unreliable test from the blocking suite immediately, so it stops eroding trust, while keeping it running non-blocking until it is fixed o…
Neighbouring topics
Testing & Quality Architecture
General material on designing a testing strategy as an architectural concern.
Test Architecture Strategy
Choosing what to verify where, given the failure modes that actually occur.
Test Pyramid Shapes
Pyramid, trophy and honeycomb, and the system properties that justify each shape.
Integration Test Boundaries
What sits inside a test's boundary, what is faked, and the confidence that follows.
Contract Testing at Scale
Keeping dozens of services compatible without an environment that runs all of them.
Consumer-Driven Contracts
Consumers declaring what they rely on, and providers verifying against those declarations.
Test Data Management
Realistic data without copying production personal data into a weaker environment.
Synthetic Data
Generating data with the shape and edge cases of the real thing, and where it misleads.
Environment Parity
The differences between staging and production that decide which bugs survive to release.
Service Virtualisation
Standing in for a dependency you cannot call, and keeping the stand-in honest.
End-to-End Test Economics
Why broad end-to-end suites get slow, flaky and abandoned, and what to keep.
Non-Functional Test Strategy
Testing availability, latency, security and recovery rather than only behaviour.
Performance Test Design
Workload models, warm-up, think time, and the distribution the average hides.
Chaos as a Test
Fault injection with a hypothesis, a blast radius and an abort condition.
Security Testing in the Pipeline
SAST, DAST, dependency and secret scanning, and what to do with the findings.
Accessibility Testing
Automated checks, their ceiling, and the manual testing that has to sit above it.
Mutation Testing
Measuring whether tests would actually notice a defect, not just cover a line.
Testing in Production
Synthetic transactions, dark launches and shadow traffic, done deliberately and safely.
Quality Gates
Thresholds that block a release, who may override them, and how they decay.