Your monorepo runs 320 pipelines a day. Median wait before a runner picks up a job is 11 minutes; the pipeline itself takes 14. Finance asks whether doubling the runner pool is worth $9k a month. Work out the number that answers them.
Show the full answer Hide the answer
The assumptions, stated
Eighty engineers push roughly four times a day, which is where 320 comes from. The question is what 11 minutes of queue wait costs, and the honest answer starts with an uncomfortable assumption: a wait shorter than the time it takes to re-establish context is close to a total loss whichever way the engineer spends it. Either they sit and watch, or they switch to something else and pay to come back. Published estimates for re-establishing context on non-trivial work cluster around 10 to 20 minutes, so for an 11-minute wait, 10 minutes of lost attention per run is a conservative figure rather than a generous one.
The arithmetic
- 320 runs per day times 10 lost minutes equals 3,200 minutes, about 53 engineer-hours a day.
- Over 21 working days that is roughly 1,100 engineer-hours a month.
- At a fully loaded $100 an hour, roughly $110k a month of attention against a $9k compute bill.
The number is a range, not a point. The assumption that dominates the error is the lost fraction, not the hourly rate: halve it and the comparison is still $55k against $9k. That is the useful property of this calculation — it is robust to being wrong by a factor of two, which is why it settles the argument in one slide.
What the number rules in
Buy down queue wait whenever runs per day times minutes saved times the loaded hourly rate exceeds the compute bill, which below about 20 minutes of wait is almost always true. The decision rule is not "is CI expensive" but "is CI more expensive than the people waiting for it".
The second-order effect matters more than the hours. Above roughly 15 minutes of feedback latency, engineers start batching work per push to amortise the wait. Bigger pushes mean bigger reviews, bigger blast radius per deploy and a higher change failure rate. The compute saving reappears as a delivery regression that nobody attributes to the runner pool.
When more runners are the wrong answer
If the wait is serialisation rather than starvation, doubling the pool changes nothing. A merge queue validating one change at a time, a licence-limited analysis step, a single shared integration environment behind a lock: in all three the jobs queue while runners sit idle.
The measurement that distinguishes the two takes an afternoon. Plot runner utilisation during the waiting periods. Starvation shows runners pinned near 100% busy with a backlog; serialisation shows idle runners and a backlog at the same moment. Only the first case is fixed by money.
What a strong answer adds
Name the third possibility: the pool is large enough on average and too small at the daily peak, because pushes cluster after stand-up and before the end of the day. That is an autoscaling problem with a floor, and the fix is cheaper than doubling the steady-state pool.
Then state what you would measure afterwards. Not runner count, but p50 and p95 queue wait, and median changed lines per push — the second confirms whether the batching effect actually reversed, which is the benefit you sold.