Two teams own the same test suite. Team A can run the relevant tests locally in 40 seconds. Team B's only way to run them is a 9-minute pipeline. Team B's pull requests are about four times larger and when a test fails they take far longer to work out which change caused it. What is the mechanism connecting feedback time to those two outcomes and which change helps most?
Show the full answer Hide the answer
The mechanism
Feedback latency sets batch size, and batch size sets the cost of finding the cause.
When a verification cycle costs 9 minutes, no rational engineer spends it on one line. They accumulate several changes and verify them together, because the cost per run is fixed and the way to amortise it is to put more into each run. So each run now verifies some number of changes k.
Then the run fails. With k changes in flight, the failure does not point at a change - it points at a batch. Narrowing it down means bisecting, roughly log2(k) further runs, each one 9 minutes. For a batch of 12 changes that is about 4 runs, so 36 minutes of waiting to answer a question that a 40-second loop answers instantly, because at 40 seconds nobody batches in the first place.
The second effect is the engineer, not the machine. A 9-minute wait is longer than the window in which someone holds a change in working memory, so they switch to something else and pay a re-entry cost on return. At 40 seconds they stay in the problem. This is why the duration matters in absolute terms and not only as a ratio: the behaviour changes somewhere around the one-to-two-minute mark, not at any particular speed-up factor.
The arithmetic is unforgiving either way. Twenty cycles in a day at 9 minutes is three hours of wall clock per engineer; at 40 seconds it is about 13 minutes.
Why a correct local subset wins
It moves the loop under the threshold where batching stops. "Correct" is the load-bearing word: the subset has to be the tests that could plausibly break given what changed, derived from a dependency or build graph rather than from a folder convention, or engineers will not trust it and will go back to waiting for the pipeline. The full suite stays in CI, where it belongs, because the local subset is a fast approximate signal and the pipeline is the authoritative one.
Why the other options fail
- More runners. Parallelism cuts queueing time and, if the suite shards, total run time. It does not shorten the critical path of a single developer's single question, which is what batch size responds to. A pipeline that starts instantly and still takes 9 minutes produces the same behaviour.
- A line-count limit on pull requests. This targets the symptom and pushes engineers to split changes that cannot be independently verified, which creates broken intermediate commits and more review overhead. Batch size is an effect of feedback cost; cap the effect and the cost just reappears elsewhere.
- 9 minutes to 6. A 33% improvement that stays on the same side of the threshold. Six minutes still exceeds working-memory span, so engineers still batch, and the bisect still costs four runs of six minutes.
- Full suite on every push. More runs of the same slow loop, more compute spend, and no change to how long an engineer waits for an answer about one change.
When this is the wrong answer
If the suite already finishes in 90 seconds, do nothing - the loop is already below the threshold and test selection is a real engineering investment, typically weeks, plus ongoing accuracy maintenance. The investment is justified when loop count per engineer-day multiplied by loop duration exceeds the total time the team spends in the outer loop, which for most teams happens well before anyone has measured it.