beginner 2 min answer Multiple choice

A team's pull-request suite has grown from 4 minutes to 26 minutes over a year, because every bug fix added a test and nothing was ever removed. No one has ever proposed deleting a test. What has the team given up, and which policy fixes it?

ci-feedback-timetest-budgetbatch-sizesuite-growthquality-gates
Pick one
Show the full answer Hide the answer

What was gained

Each of those tests was added for a real reason - a defect that reached production once. Individually every addition was correct, and that is why nobody stopped it.

What was paid

Not compute. Merge frequency, and through it batch size. The costs land at two thresholds that have nothing to do with the CI bill:

  • Around two minutes, developers stop running the suite locally before pushing. First feedback moves from the editor to CI, so every mistake now costs a push and a wait instead of a re-run.
  • Around ten minutes, developers stop waiting for the result and switch tasks. A red build now costs a context reload rather than a fix, and it arrives after the author's attention has moved.

At 26 minutes with eight engineers merging six times a day, the pipeline serialises roughly two and a half hours of the working day, and the rational individual response is to batch more work into each change. Larger changes are harder to review, harder to bisect and riskier to roll back, so a slow gate degrades every downstream safety property while every individual test that caused it was justified.

The policy

Give each stage a fixed budget - roughly two minutes for anything run pre-commit, ten for the pull-request gate, an hour for the full pipeline - and fail the build when the stage exceeds it, the same way it fails on a broken test. Pair it with a displacement rule: a new test that pushes a stage over budget ships with either a deletion or a move to a slower stage.

The mechanism that matters is social, not technical. A fixed budget makes tests compete, which forces someone to state what each one is worth. Without scarcity, no test is ever removed, because removing a test has a visible downside and no visible upside.

Why the other options fail

  • Run the suite only after merge. Feedback moves from before the mistake is shared to after, and main is now red for everyone. This trades a delay for a coordination failure, which is a worse currency.
  • Parallelise until it fits. The right answer when tests are genuinely independent and a worker costs less than an engineer-hour, and it should probably be done anyway. It does nothing about the arrival rate of new tests, so it buys back the window once and postpones this exact conversation by about a year. It also masks a suite where 40% of the tests are duplicates.
  • Mark slow tests optional. An optional test is a test nobody runs and everybody still maintains. This is the worst outcome: the cost stays and the signal goes.

When the budget is the wrong tool

A team of three on a pre-launch product should ignore this and write tests, because the batch-size cost only bites when several people merge into the same branch on the same day. And a nightly or release-candidate stage has no interactive user waiting, so it should be budgeted on infrastructure cost and on how long a release can be blocked, not on human attention.