beginner 3 min answer

A service has 95% line coverage and 3,000 passing tests. A bug shipped that made every order for one country total to zero. Why did the suite not catch it, and what kind of test would have?

coveragetest designassertionsbeginnerquality
Show the full answer Hide the answer

The mechanism

Line coverage measures which lines ran, not whether anything checked the result. A test that calls calculateTotal(order) and asserts only that it did not throw will mark every line of that function covered. The suite is reporting how much of the code was executed during testing, and that is a different question from how much of it was verified.

So 95% coverage is consistent with a suite that would pass if the function returned zero for everything, provided no assertion compares the number to an expected value.

The second gap is narrower and more common than the first: coverage says nothing about which inputs were used. The tests almost certainly covered calculateTotal with a domestic order. The bug is in the branch taken for one country — a tax rule, a currency, a rounding mode — and that branch either was not exercised at all, or was exercised with an assertion loose enough not to notice.

The consequence people miss

Coverage is bounded by what the test author imagined. Every test encodes a case someone thought of. A suite with 3,000 tests represents 3,000 anticipated situations, and the bug was in the 3,001st. Raising coverage to 100% would not have helped, because the uncovered 5% is usually error handling and logging, not the case nobody considered.

This is why "we need more coverage" is almost always the wrong response to an escaped bug. The right response starts with: what class of input was never tried?

What would have caught it

In rough order of what gives most per unit of effort:

  • A test with real data from each market, not one synthetic order. Table-driven tests over a representative set of countries, currencies and tax regimes turn a single imagined case into dozens. This is the cheapest fix available and the one most teams skip.
  • Property-based tests. Instead of "a €20 order plus 20% tax is €24", assert the invariant: the total is never zero for a non-empty order; the total is never less than the sum of line items. The framework generates hundreds of inputs including ones nobody would write, and the invariant catches the whole class rather than one instance.
  • A business-invariant check in production. A continuous assertion that no completed order totals zero, alerting within minutes. Testing cannot enumerate every case, so something must watch the real ones.
  • Mutation testing, to audit the suite itself. It changes the code deliberately — flip a comparison, return zero — and reports which mutants the tests failed to kill. A surviving "return 0" mutant in calculateTotal is a direct, mechanical statement that no test checks the value.

When coverage is still worth measuring

Coverage is a useful negative signal and a worthless positive one. Zero coverage on a module means nobody has tested it, which is actionable. Ninety-five percent coverage means nothing in particular. Keep it as a floor that stops obviously untested code merging, set the floor low enough to be honest, and stop treating increases in it as progress.

Common weak answers

  • "Raise the coverage target to 100%." The missing 5% is error paths. The bug was in covered code.
  • "Add an end-to-end test for checkout." One more imagined case, at the highest possible cost per case, and it would have used a domestic order like all the others.
  • "The QA team should have caught it." Relocates the question. Whoever tests it manually is also working from imagined cases, and more slowly.
  • "Write a regression test for this bug." Correct and necessary, and it protects against this exact bug recurring, which is the least likely bug to recur. It does nothing about the class.