metric

Deploy Batch Size

also called Release Batch Size, Change Batch Size

The number of independent changes in one release, which sets the size of the candidate set during an incident and therefore dominates the time to attribute a regression and the willingness to revert.

etsydeploymentbatch sizeattributionincident response

A release ships every three weeks with about 60 merged changes in it. A conversion metric drops the next morning. Nobody knows which change did it, and the incident channel is not discussing the system at all: it is discussing whose feature gets withdrawn if the release is rolled back.

Batch size decides both of those conversations. Bisecting 60 candidates takes about 6 steps, and at roughly 25 minutes of build and deploy plus 30 minutes of observation per step that is about 5.5 hours with the regression live. At a batch size of 1 the deploy annotation on the graph is the answer.

Why it matters

Mean time to repair is dominated by time to diagnose, and time to diagnose scales with the size of the candidate set rather than the complexity of the system. Teams attack repair time with better dashboards, which help, and ignore the variable that multiplies every diagnosis.

The second effect is larger. A large batch makes revert politically expensive, because reverting 60 changes withdraws 59 innocent ones, so the team debugs forward under pressure with the fault still live. That negotiation is frequently the longest single interval in the timeline.

Implementation patterns

  • Target one to five changes per release, which requires a pipeline shorter than the interval between merges, since a 90-minute pipeline cannot support ten deploys a day.
  • One-command revert with no schema decision in it. Expand-and-contract migrations exist so that revert stays available: ship the schema change separately and ahead.
  • Annotate every graph with deploys. The annotation is what converts a small candidate set into free attribution.
  • Decouple deploy from release using flags, so the deployed batch stays small even when the business wants features to appear together.
  • Instrument the metric: changes per release, releases per day, and the age of the oldest unreleased merge. The third number predicts the next bad night.

Industry example

Etsy is the documented reference point. It released StatsD in 2011 alongside a post titled "Measure Anything, Measure Everything", annotated its graphs per deploy, and its engineers reported production deploy rates on the order of 25 a day around 2011 to 2012. The frequency is the headline that gets repeated. The mechanism that makes it safe is the small candidate set behind every movement on a graph, plus flags and staged ramps (staff first, then a small percentage of users) so exposure is controlled independently of deployment. The practice is not "deploy often because it is modern", it is "keep the candidate set small so attribution and revert stay cheap".

Failure scenarios

  • The five-hour bisect, during which the regression is live and customer-visible.
  • The revert nobody will authorise, so the team debugs forward and the incident runs into a second day.
  • Coupled migrations. A schema change inside the batch makes revert unsafe, converting a reversible decision into a one-way door discovered mid-incident.
  • Batch size cut without observability. Ten deploys a day with no per-deploy annotation gives ten unexplained graph movements instead of one.
  • The slow-signal trap. Bisecting a metric that needs a day of data to be significant, so every step looks clean.

Trade-offs

Choose Gains Pays
Small batches Near-free attribution; cheap revert; shorter incidents Pipeline investment; migration discipline; more release events to observe
Large batches Fewer release events; coordinated feature launches; less pipeline work Expensive attribution; politically blocked reverts; longer incidents

Small batches change the shape of the failure distribution as well as its size: more frequent, smaller, better-attributed incidents instead of fewer, larger, mysterious ones. An organisation that measures incident count rather than impact reads that as a regression.

When not to use it

Where the deploy itself is expensive and risky rather than cheap, batch size is not the lever. Firmware, an app-store binary, a change inside a regulated window, or a database engine upgrade all have a fixed high cost per release, and shrinking batches multiplies that cost without improving attribution. There, invest in exposure control instead: ship the code dark, ramp by cohort, and keep the ability to turn one change off without shipping anything. The same applies when the signal of interest takes a day to become significant, because bisection cannot work faster than the metric.

Interview question

Q: A team releases every three weeks with roughly 60 changes per release and wants to cut mean time to repair. They propose better dashboards. What do you propose instead, and how would you justify it with numbers?

What a strong answer covers: the bisection arithmetic at batch size 60 against batch size 1; the revert negotiation as the longest interval and why batch size causes it; the prerequisites (pipeline duration, expand-and-contract migrations, deploy annotations) rather than just asking for more deploys; and the slow-signal case, where exposure control replaces batch reduction.

Quick check

Quiz: Why does a batch of 60 changes make revert expensive even when it is technically easy? Answer: because it withdraws 59 innocent changes, so the decision needs agreement from several owners, and that negotiation usually outlasts the diagnosis.

Flashcard: What dominates time to diagnose a regression? The size of the candidate set: bisecting 60 changes costs about 6 deploy-and-observe cycles, roughly 5.5 hours, while a batch of one makes the deploy annotation the answer.