Zoom's CEO wrote on 1 April 2020 that maximum daily meeting participants went from roughly 10 million at the end of December 2019 to more than 200 million in March 2020. For any platform absorbing a 20× demand rise over three months rather than three minutes, which capacity decision does the slower surge force that the fast one does not?
Show the full answer Hide the answer
The situation, and only what is documented
The published facts are the demand figures: roughly 10 million maximum daily meeting participants at the end of December 2019, more than 200 million in March 2020, and a figure of roughly 300 million by late April. Zoom's engineering response is not public at that level of detail, so what follows is the capacity logic the shape of that curve imposes on anyone, not an account of what Zoom built.
What the shape of the curve forces
Compound the growth and the problem states itself. Going 20× in 90 days is about 3.4% per day, a doubling roughly every 21 days. If any input you need takes 45 days to obtain, you are permanently two doublings behind — the capacity that arrives was sized for a quarter of the load now present. That single comparison, growth rate against lead time, is what separates a three-month surge from a three-minute spike.
The inputs with lead times are the ones nobody models: physical racks and the transit and peering contracts that feed them, cloud quota increases that pass through human approval, address space, vendor licence tiers priced per seat, specialised hardware with manufacturing queues, and people who can operate what you bought. Autoscaling cannot allocate capacity that does not exist or is not permitted, and a quota denial looks exactly like a capacity shortage from inside the service.
A 20× rise also crosses architectural thresholds the model did not contain. A control plane sized for 10 million participants, a configuration store, an identity service, a metrics pipeline whose cost is linear in sessions — each has a point where the next factor of two is a redesign rather than a purchase.
The decision rule
When the demand growth rate exceeds one over the longest lead time, stop forecasting and start buying optionality. In practice: standing quota headroom you never use, pre-approved secondary regions, contracts with uplift clauses rather than fixed commitments, and a degradation ladder that is already shipped and tested.
When this is the wrong lesson to copy
Zoom's workload shards naturally — a meeting is an independent unit of media routing, and adding capacity adds meetings. A product whose load converges on one transactional core gets no such gift. There, a 20× surge lands on a single write path and no amount of quota, region or contract fixes it; the work is partitioning, which takes quarters.
The second mistake is building for this at all. Almost no product sees 20× in a quarter, and the capital held against a forecast nobody believes is money removed from the product. The defensible position is cheap optionality plus a rehearsed degradation path, not provisioned capacity for a surge that will not come.
Why the other options fail
- Pre-warming pools and caches is the correct answer to the three-minute spike, where the constraint is a control loop that cannot react fast enough. Over three months every pool warms itself many times over, and pre-warming addresses nothing that binds.
- Raising the autoscaling maximum and shortening the cooldown is the most common wrong answer because it feels like capacity work. It is the right move when the limit is reactive scaling speed. Here the scaler is already waiting on quota, hardware or address space, and a higher maximum simply fails faster.
- Shedding at admission with a visible queue is genuinely valuable and belongs in the degradation ladder, but it is what you do when capacity is structurally short. It does not change what capacity you can obtain, and treating it as the answer concedes the growth.
- Shortening health-check intervals improves replacement time after a failure, which is a resilience lever and not a capacity one. At 20× growth the instances are not failing; there are not enough of them.