A service has 8,000 unit tests running in 4 minutes, over about 60,000 lines of code. Someone proposes mutation testing in CI. Estimate the runtime, and say what that number means for where it belongs.
Show the full answer Hide the answer
The assumptions, stated
- Mutants generated. Mutation operators (flip a comparison, change a boundary, return a default, remove a call) yield roughly one mutant per 2–5 lines of mutable code. Not all 60,000 lines are mutable — imports, declarations, generated code. Take 30,000 mutable lines and one mutant per 3, so about 10,000 mutants.
- Cost per mutant. Naively, each mutant requires a test run: 10,000 × 4 minutes = 40,000 minutes, or roughly 28 days. That number is the reason nobody runs mutation testing naively.
- The optimisation that matters. A mutant only needs the tests that cover the mutated line. With coverage data, a mutant typically runs a few dozen tests rather than 8,000 — call it 0.5% of the suite, about 1.2 seconds — plus JVM or interpreter startup, which modern tools amortise by reusing a warm process.
The arithmetic
10,000 mutants × 1.2 s ≈ 3.3 hours single-threaded
÷ 16 parallel workers ≈ 12 minutes
So the honest range is 10 minutes to about an hour on a well-parallelised build, depending on how concentrated coverage is. Add a factor of 2–3 if the tests are not pure units — anything touching a database, the filesystem or the network destroys the per-mutant cost model, and that is the assumption most likely to be wrong.
Which assumption dominates the error
Tests per mutant, by a wide margin. The model assumes a mutated line is covered by a handful of focused tests. If the suite is mostly broad integration tests that each touch thousands of lines, every mutant runs a large fraction of the suite and the estimate rises by one to two orders of magnitude — straight back towards the 28-day figure.
That sensitivity is itself informative: mutation testing is cheap exactly when the suite is well-factored, and ruinous when it is not. The runtime estimate doubles as a measurement of test design.
What the number rules in and out
Ruled out: the pull-request gate. Even 12 minutes on top of a 4-minute suite is a tripling of the feedback loop, and the output is not a pass/fail that a developer can act on in the moment — it is a list of surviving mutants requiring judgement.
Ruled in, two ways:
- Incremental mutation testing on the diff. Mutate only lines the pull request changed. A 200-line change produces perhaps 70 mutants, runs in well under a minute, and gives feedback precisely where attention already is. This is the form that survives contact with a real team, and it is how the practice is viable at all in CI.
- Full runs nightly or weekly, reported as a trend with surviving mutants triaged into the backlog. Not a gate.
What it is actually for
Not a score. A mutation score of 70% versus 75% is not a meaningful distinction and chasing it produces tests written to kill mutants rather than to express intent. The value is in the specific survivors: a mutant that changed >= to > on the refund-eligibility boundary survived, meaning no test checks that boundary. That is a precise, actionable statement about one line, and it is the kind of statement no coverage report can make.
When not to run it at all
It is the wrong investment for a suite with weak assertions. Mutation testing audits how well your tests detect changes, which presupposes tests that assert on behaviour. If the suite is full of tests that call a function and check it did not throw, mutation testing will report a catastrophic score and tell you nothing you could not have learned by reading three test files. Fix the assertions first; the tool's output is only useful once it is mostly green.
It also has little to offer code whose correctness is about integration rather than logic — glue, configuration, orchestration. Mutating a line that wires two components together usually produces either an unkillable mutant or a trivially killed one. Point it at the modules where the logic is: pricing, eligibility, tax, permissions, state machines. A 400-line pricing engine is worth more mutation attention than the other 59,600 lines combined.