advanced 4 min answer

"Our test pyramid is upside down — mostly end-to-end tests, few unit tests. How would you fix it?" Treat this as the interview prompt it is.

test pyramidtest strategyinterviewtrade-offsfeedback loop
Show the full answer Hide the answer

What the interviewer is testing

Whether you accept the framing. The weak answer takes "the pyramid is upside down" as the problem statement and proposes a plan to invert it. The strong answer asks what the shape is costing, because an inverted pyramid is a symptom and occasionally not even a defect.

It is also a test of whether you reason from properties or from diagrams. The pyramid is a heuristic about feedback speed and failure diagnosability, not a target ratio, and a candidate who treats it as a ratio to hit has memorised the picture.

The clarifying questions that change the answer

  • What is the suite costing you right now? Runtime, flake rate, time to diagnose a failure, and how often people re-run rather than investigate. If the E2E suite runs in six minutes and fails only for real reasons, there is no problem and the right answer is to leave it alone and say so.
  • What is the architecture? The pyramid assumes a system whose logic lives in units worth testing. A thin service that mostly orchestrates calls to other services has almost no unit-testable logic, and its correct test shape genuinely is integration-heavy. Prescribing a pyramid there produces tests of mocks calling mocks.
  • Where do production escapes actually come from? The only evidence that matters. If escapes are integration and configuration failures, more unit tests will not reduce them, whatever the shape improves to.
  • Is it one deployable or twelve? A modular monolith can test realistically and fast in-process. A distributed system cannot, and its E2E tests carry a cost that no amount of discipline removes.
  • How long has the team got? Inverting a pyramid is quarters. A plan that does not sequence value early will not survive contact with delivery pressure.

A strong answer's arc

Diagnose, then convert, then prune — in that order, and never by deleting first.

  1. Measure before changing. Runtime, flake rate, failure-diagnosis time, and escape categories from the last twenty incidents. This takes a day and it determines everything after.
  2. Attack flakiness before shape. A flaky E2E suite is a worse problem than a top-heavy one, and fixing it is faster. An inverted pyramid that is reliable is a tolerable situation; any pyramid that is flaky is not.
  3. Convert, do not delete. For each slow E2E test, find the logic it actually verifies and write that as a unit or component test; only then remove the E2E case. Deleting first trades a slow suite for no suite, and it is how this initiative destroys trust in one sprint.
  4. Add the middle layer the shape is missing. Most inverted pyramids lack component tests — a service tested in-process with its real logic and stubbed boundaries. This layer gives most of E2E's confidence at close to unit speed, and it is usually where the real win is.
  5. Target an E2E suite that is deliberately small and deliberately chosen: the handful of critical user journeys, run against a real deployment, with everything else pushed down. Ten to thirty, not a thousand.
  6. Stop the regrowth. A budget on E2E suite runtime, enforced in CI. Without it the shape re-inverts within a year, because adding an E2E test is always the locally cheapest way to cover a new case.

Common weak answers

  • Proposing a ratio. "70/20/10" is a description of some healthy suites, not a goal. Hitting it by writing low-value unit tests for getters is worse than the starting position.
  • Deleting E2E tests to improve the shape. Fastest way to an incident and the fastest way to lose the mandate.
  • Ignoring flakiness because it is not a shape problem. It is the problem that makes every other number meaningless.
  • Applying the pyramid uniformly across services. The right shape for a pricing engine and for an API gateway are different, and a single mandate produces ceremony in the orchestration services.

What a strong answer adds

Three things at the senior level:

  • Connecting suite shape to delivery metrics. Feedback time is lead time. A 40-minute suite caps how often anyone can safely merge, which sets deploy frequency, which sets batch size, which sets incident severity. That chain is how you get funding for test work, and it is the argument that works on people who do not care about pyramids.
  • Naming when the pyramid does not apply. Thin orchestration services, data pipelines where the logic is in SQL, ML systems where correctness is statistical. A candidate who can say where the heuristic stops is one who understands it.
  • Treating the test suite as a product with an owner and a budget. Suites decay by default because adding is always cheaper than curating, and no shape survives without someone whose job includes saying no.