beginner 3 min answer

Every one of 180 services must move to a new authentication library. The platform team migrated three sample services and each took about two days. A director asks for the total. Roughly what is it, and which assumption decides the answer?

estimationmigrationlong tailbeginnerfleet
Show the full answer Hide the answer

The assumptions, stated

180 services. A median of two engineer-days, from three samples. And the assumption nobody writes down: the platform team chose those three samples, which means they picked the ones they understood. A hand-picked sample tells you about the easy end of the distribution and nothing about its shape.

The arithmetic

The number that gets reported is 180 × 2 = 360 engineer-days, about 1.6 engineer-years at 220 working days. It is wrong because it uses the median as though the distribution were narrow.

Fleet work is long-tailed, and the tail has recognisable causes: a service with no owner, a forked copy of the old library, a hand-rolled auth path, a framework two majors behind, or no test suite to prove the change is safe. Ten to twenty per cent of a mature estate is usually in one of those states. Take 15% at fifteen days: 27 × 15 = 405 engineer-days — more than the entire median term.

Total: roughly 765 engineer-days, about 3.5 engineer-years, with a defensible range of 2.5 to 5.

Which assumption dominates the error

Not the median, and not the count. The tail share and the tail cost. Moving the median from two days to three adds 180 days. Moving the tail from 15% to 25% adds about 270. So the cheapest way to tighten this estimate is to measure the tail, not to re-measure the easy case.

That measurement is a one-day scripted inventory run against what is actually in production: for each of the 180 services, does it have a named owner, does its test suite run, and is its library version the one the samples used. Three columns turn a factor-of-three guess into roughly a factor of 1.5, and the script is reusable for the next fleet change.

What the number rules in and out

3.5 engineer-years spread over 26 teams is 1 to 2% of their year, which sounds harmless and is exactly why fleet mandates fail: the average is trivial and the tail is not, and the tail lands on whichever teams are least able to absorb it. It rules out "everyone does this in their next sprint", because the 27 hard services are not in the teams that have slack. It rules in a different shape of plan: a scripted change plus review for the 153 straightforward services, and a dedicated pair who do nothing but the tail — 27 services at 15 days is about a quarter's work for two people, and it is more than half the total effort.

Choose the mandate only when the tail term is smaller than the median term; above that, fund a migration team, because a mandate distributes the hard cases to the people with the least context and no deadline they own.

Common weak answers

  • Median times count, reported as a single number. It is the one answer that is reliably wrong by a factor of two or more.
  • A flat 30% contingency. Contingency hides the structure, cannot be challenged, and is always too small when the tail is this heavy.
  • Refusing to give a number. The director needs a figure to decide whether this is a mandate or a funded project, and those two answers differ by an order of magnitude in organisational cost.

When this is the wrong answer

Where the services really are uniform — generated from one template, same version, one owner, tests everywhere — the median is the estimate and the inventory is wasted effort. The test is cheap: look at the variance in your samples. Two days, two days and two days from a hand-picked set tells you nothing; two days, two days and nine days tells you the tail exists and you should go and size it.