Mutation Testing
Measuring whether tests would actually notice a defect, not just cover a line.
3 to work through
-
advanced
A team adopts mutation testing to stop shipping defects in code that coverage says is tested. On a 40,000-line service with a 4-minute unit suite, what have they bought, what are they paying, and when does that bill arrive?
3 min answer -
advanced
A team has 85% line coverage and keeps shipping defects in tested code. What would reveal the gap?
2 min answer -
advanced
What does mutation testing reveal that coverage does not, and where is it worth the cost?
2 min answer
4 terms in this topic
Assertion Strength
Whether a test would actually fail if the code were wrong - which coverage does not measure and mutation testing does, and which explains defects in …
practiceMutation Scope Selection
Restricting mutation analysis to the lines a change touches and to code whose failure is expensive, so a technique that costs one to two orders of ma…
metricMutation Score
The proportion of deliberately introduced faults that the test suite detects — a measure of whether tests would notice a defect, unlike coverage.
toolMutation Testing
Deliberately introducing small faults into the code and measuring how many the test suite detects, as a measure of test strength rather than test reach.
Neighbouring topics
Testing & Quality Architecture
General material on designing a testing strategy as an architectural concern.
Test Architecture Strategy
Choosing what to verify where, given the failure modes that actually occur.
Test Pyramid Shapes
Pyramid, trophy and honeycomb, and the system properties that justify each shape.
Integration Test Boundaries
What sits inside a test's boundary, what is faked, and the confidence that follows.
Contract Testing at Scale
Keeping dozens of services compatible without an environment that runs all of them.
Consumer-Driven Contracts
Consumers declaring what they rely on, and providers verifying against those declarations.
Test Data Management
Realistic data without copying production personal data into a weaker environment.
Synthetic Data
Generating data with the shape and edge cases of the real thing, and where it misleads.
Environment Parity
The differences between staging and production that decide which bugs survive to release.
Service Virtualisation
Standing in for a dependency you cannot call, and keeping the stand-in honest.
End-to-End Test Economics
Why broad end-to-end suites get slow, flaky and abandoned, and what to keep.
Non-Functional Test Strategy
Testing availability, latency, security and recovery rather than only behaviour.
Performance Test Design
Workload models, warm-up, think time, and the distribution the average hides.
Chaos as a Test
Fault injection with a hypothesis, a blast radius and an abort condition.
Security Testing in the Pipeline
SAST, DAST, dependency and secret scanning, and what to do with the findings.
Accessibility Testing
Automated checks, their ceiling, and the manual testing that has to sit above it.
Flaky Test Management
Quarantine, detection, and the trust a suite loses once red stops meaning broken.
Testing in Production
Synthetic transactions, dark launches and shadow traffic, done deliberately and safely.
Quality Gates
Thresholds that block a release, who may override them, and how they decay.