concept

Capacity Lead Time

also called Provisioning Lead Time

The delay between deciding capacity is needed and having it available, which determines whether a capacity model can forecast at all or must instead buy optionality.

capacity-modellingquotaprocurementsurgeoptionality

A capacity model is usually judged on the accuracy of its forecast. The more useful test is a comparison: does the forecast horizon exceed the lead time of the capacity it is forecasting? If it does not, the model cannot converge however good the forecast is, because the capacity arrives after it was needed.

Lead time is not one number. It is the longest of several, and the longest one is rarely compute.

Why it matters

Compound growth makes the comparison brutal. Zoom's CEO wrote on 1 April 2020 that maximum daily meeting participants went from roughly 10 million at the end of December 2019 to more than 200 million in March 2020. That is roughly 3.4% per day — a doubling about every 21 days. Against a 45-day lead time you are permanently two doublings behind: whatever arrives was sized for a quarter of the load already present.

The inputs with lead times are the ones models omit: racks and the transit and peering contracts that feed them, cloud quota increases that pass through human approval, address space, per-seat licence tiers, specialised hardware with manufacturing queues, and people who can operate what you bought. Autoscaling allocates only capacity that already exists and is already permitted, and from inside a service a quota denial is indistinguishable from a capacity shortage.

Implementation patterns

  • Inventory the lead times. One table: resource, lead time, who approves, and what happens at the limit. The act of filling it in usually finds a 30-day item nobody knew about.
  • Compare growth rate to one over the longest lead time. Above that line, stop refining the forecast and buy optionality.
  • Hold standing quota headroom — commonly two to three times current peak — and audit it quarterly, because quota is free to hold and slow to obtain.
  • Pre-approve a secondary region or account and run a token workload in it, so capacity there is a scaling decision rather than a project.
  • Negotiate uplift clauses rather than fixed commitments, so a contract does not become the ceiling.
  • Ship and rehearse a degradation ladder, because the only capacity available instantly is the capacity you stop consuming.

Industry example

Zoom's published figures are the growth, not the engineering response, and the capacity logic they impose is general: a 20× rise over three months is a procurement and permission problem, where a 20× rise over three minutes is a control-loop problem. The two surges have almost disjoint remedies — pre-warmed pools and pre-scaled fleets for the fast one, quota and contracts and optionality for the slow one — and teams routinely apply the fast remedy to the slow problem because it is the one engineering controls.

Failure scenarios

  • A quota denial during a surge, which surfaces as capacity errors and takes days of support correspondence to clear.
  • Hardware ordered against a forecast that moved, arriving sized for a load two doublings old.
  • A per-account or per-region limit discovered at the limit, where the workaround is a new account and a migration.
  • No one to operate it: capacity delivered into a team that cannot absorb the operational load, so the capacity sits unused.
  • An architectural threshold mistaken for a capacity shortage, where the next factor of two needs a redesign of a control plane or a single write path and no amount of provisioning helps.

Trade-offs

Optionality costs money for capacity you may never use: reserved commitments you do not consume, a secondary region kept warm, contracts priced for an uplift that does not come. It is a small continuous cost.

Provisioning to a forecast costs capital held against a number nobody believes, and the capital is removed from the product. The defensible position for almost every organisation is cheap optionality plus a rehearsed degradation path, not provisioned capacity for a surge that probably will not arrive — and the rarity of 20× quarters is the reason.

When not to use it

Where every input is elastic and the growth is slow, the comparison is trivially satisfied and the inventory is busywork. Do the exercise when growth is fast, when hardware or contracts are in the path, or when a regulator or a data-residency rule makes a region a project rather than a configuration value.

Interview question

Q: "Demand is doubling every three weeks and your longest capacity lead time is six weeks. Leadership asks for a 12-month capacity plan. What do you give them instead, and how do you justify it?"

What a strong answer covers: the growth-rate against lead-time comparison and why forecasting cannot converge; the inventory of non-compute lead times including quota and contracts and people; optionality as the substitute for a forecast; a shipped degradation ladder as instantly available capacity; and the honest cost of optionality against the cost of being two doublings behind.

Quick check

Quiz: Growth doubles every 21 days and lead time is 45 days. What does that imply? You are permanently about two doublings behind, so the model must buy optionality rather than forecast.

Flashcard: Why does autoscaling not solve a three-month 20× surge? — It allocates only capacity that already exists and is permitted, and the binding constraints are quota, contracts, hardware and people.