beginner 3 min answer Multiple choice

A team has 3200 unit tests that run in 90 seconds, 40 integration tests and 12 end-to-end tests. Over six months it shipped 19 defects to production and 14 of them were wiring faults - a wrong column name in a query, a header the downstream service rejected, a queue bound to the wrong exchange. Which change to the suite's shape actually addresses that?

test-pyramiddefect-escapeintegration-testswiring-defectssuite-shape
Pick one
Show the full answer Hide the answer

What is being tested

Whether you reshape a test suite from evidence or from a diagram. The pyramid is an argument about cost per test, not about where risk lives. This team's risk is measurably at the seams, and no amount of the cheap layer touches it.

The reasoning

Every one of the 14 escapes shares a property: the unit under test was correct and the connection to something outside the process was not. A unit test substitutes that outside thing, so by construction it cannot observe a wrong column name, a rejected header or a misbound exchange. The 3200 tests are not weak. They are pointed somewhere the defects are not.

The diagnostic that produced this answer is cheap and almost nobody runs it: take the last twenty production defects and label each with the cheapest layer that could have caught it. Not the layer that would have been nicest - the cheapest one that would have gone red. Here the label is "service started against a real database and a real broker" fourteen times out of nineteen. That histogram is the suite's target shape, and it is specific to this codebase.

The decision rule: grow the layer where your escape histogram has its mode, and hold every other layer at its current size until the histogram moves. It flips when the histogram flips - a team whose escapes are arithmetic errors in pricing logic should be doing the opposite, because those defects are cheapest to catch in-process and an integration test that finds one costs fifty times as much to run and diagnose.

What it costs: integration tests against real infrastructure run in seconds rather than milliseconds, need container lifecycle management, and have a wider diagnostic radius. A realistic budget is a few hundred of them at five to fifteen seconds each, parallelised, which is minutes rather than the current 90 seconds. That is the price of moving the suite to where the defects are.

Why the other options fail

  • Raise the coverage threshold from 70% to 90%. Coverage measures which lines executed during the suite, and all 14 defects live in code that already executes. The extra 20% will be written against the easiest remaining lines, because that is what a threshold rewards. It buys test-writing effort and no new defect class.
  • Add end-to-end tests for each journey where the defects appeared. This catches the same faults at the highest possible price: slowest to run, widest set of causes consistent with a red result, and most likely to be flaky. It also fixes the past rather than the class - the next wiring fault will be on a journey nobody listed.
  • Introduce mutation testing on the existing unit tests. Mutation testing measures whether tests would notice a fault in the code they cover. These tests cover the right code and the faults are outside it, so a perfect mutation score changes nothing here.

When this is the wrong answer

If the escape histogram is flat, the problem is not shape - it is that the team ships changes nobody reviewed against requirements nobody wrote, and reshaping the suite treats a symptom. And a service that is genuinely pure logic with one datastore gets very little from growing this layer, because there is only one seam and a handful of tests cover it.

What a strong answer adds

Re-run the histogram every quarter, because the shape that is right today is a property of the current architecture. And say the uncomfortable part out loud: the 3200 unit tests are probably too many, and the displacement rule - a new test in the slow layer must be paid for by a removal - is the only force that ever shrinks a suite.