A team ships one release every three weeks containing about 60 merged changes. A conversion regression appears after the latest release and nobody knows which change caused it. Estimate the time to attribute the regression at a batch size of 60 against a batch size of 1, and say which number to target first.
Show the full answer Hide the answer
The assumptions, stated
60 candidate changes. A build-and-deploy cycle takes about 25 minutes. The regression is visible in a metric after roughly 30 minutes of traffic. Assume for now that it reproduces deterministically and that one change is responsible.
The arithmetic
Bisecting 60 candidates takes log2(60), about 6 steps. Each step costs 25 minutes of pipeline plus 30 minutes of observation, so roughly 55 minutes a step and about 5.5 hours of elapsed time, all of it with the regression live. A linear walk backwards through the 60 changes, which is what teams actually do when the changes are not independently deployable, costs up to 55 hours.
At a batch size of 1 the candidate set is 1. Attribution costs nothing: the deploy annotation on the graph is the answer, and the elapsed time collapses to the time to notice.
The second number matters more than the first. Reverting a release of 60 changes reverts 59 innocent ones, so instead of reverting, the team negotiates about which product owners lose their feature tonight. That negotiation, not the diagnosis, is usually the longest interval in the incident. At a batch size of 1, revert is a one-line decision nobody argues with.
Etsy is the documented reference point here: it released StatsD in 2011 with a post titled "Measure Anything, Measure Everything", annotated its graphs per deploy, and its own engineers reported production deploy rates on the order of 25 a day around 2011 to 2012. The deploy rate is the headline; the mechanism that makes it safe is the small candidate set behind every graph movement.
Which assumption dominates the error
The observation window, not the pipeline. If the regression is only visible over 24 hours of data - true for ranking, pricing and most conversion metrics - then bisection takes a month at any batch size and the arithmetic above is useless. That case needs exposure control rather than smaller batches: a flag per change, a ramp, and a cohort comparison that gives a signal per change simultaneously rather than one change at a time.
What the number rules in
Target a batch size of one to five changes, which requires: a pipeline shorter than the interval between merges, a revert that is one command and does not need a schema decision, and one observable metric per change. Note what is not on that list: a new architecture. Batch size is a property of the release process, and most of the cost of fixing it sits in test runtime and database migration discipline, not in service boundaries.
When this is the wrong answer
Where the deploy itself is expensive and risky rather than cheap - firmware, an app-store binary, a change inside a regulated window - the batch is fixed by something you do not control. Then stop optimising batch size and optimise exposure: ship the code dark, ramp by cohort, and keep the ability to turn one change off without shipping anything.