Pre-Scaling
also called Scheduled Scaling, Predictive Scaling
Provisioning capacity before demand arrives, based on a forecast or a known event, because autoscaling cannot respond faster than its provisioning loop.
Autoscaling is a closed loop: observe a metric, decide, provision, warm, admit to rotation. Every stage has latency, and the total is typically minutes. Any demand change arriving faster than that loop is met by whatever capacity already exists.
Pre-scaling accepts this and moves the decision earlier — from a reactive metric to a forecast or a calendar.
Why it matters
The most damaging capacity incidents are not gradual overloads; they are step functions. A scheduled live event, a product launch, a flash sale, a marketing send, an election result, a viral moment. In every case, by the time the metric moves, the outcome is already determined by the capacity in place.
Worse, the shortfall is self-amplifying: latency rises, clients time out and retry, and the added load can push the system into a state that new capacity no longer resolves, because a large share of the traffic is now retries.
Implementation patterns
- Scheduled scaling from an event calendar. This is more organisational than technical: the platform must know what is scheduled, which requires a path from marketing, product and partnerships into capacity planning.
- Leading indicators rather than lagging ones. Concurrent connections, request arrival rate and queue depth move before CPU does. Better still, domain signals — stream starts, sessions created, checkout initiations — predict load rather than reflecting it.
- Warm pools. Instances provisioned and initialised, held out of rotation, added in seconds rather than minutes. Costs idle capacity; buys a response time autoscaling cannot provide.
- Asymmetric policies. Scale up fast and aggressively; scale down slowly and conservatively. An hour of excess capacity is trivial; removing capacity before a second spike is an outage.
- Static stability during the peak. Nothing that must succeed during the spike — a scaling API call, an image pull, a configuration fetch — should be on the critical path, because anything required during the spike can fail during the spike.
Industry example
A live-streaming platform expecting a tenfold viewer increase within minutes cannot use CPU-based autoscaling, for two independent reasons. First, timing: metric delay plus evaluation interval plus stabilisation plus provisioning plus warming exceeds the entire event ramp. Second, signal: a server holding many idle connections has low CPU while approaching its memory and file-descriptor limits, so CPU is measuring the wrong axis entirely.
The design that works pre-scales from the event schedule, scales on connection count and viewer-join rate, holds warm pools, and treats autoscaling as a background correction. Admission control is the backstop: when capacity genuinely runs out, cap new joins and degrade bitrate deliberately rather than collapsing.
The same reasoning governs commerce platforms before a flash sale, notification platforms before a broadcast send, and news platforms before an election night.
Failure scenarios
- Treating autoscaling as burst protection, which is the single most common capacity design error.
- No calendar integration, so engineering learns about the event from the graph.
- Forecast without a floor. A forecast that is wrong low is an outage; pre-scaling needs a margin and a degradation path for when the margin is wrong.
- Scaling down immediately after the peak, into the second wave.
- Warm pools that are not actually warm — instances started but with cold caches and empty connection pools, which fail on first real traffic.
Trade-offs
Pre-scaling costs money for capacity that is idle, sometimes for an event that does not materialise. It also requires forecasting, which is work, and organisational plumbing, which is harder.
Against that, it is the only mechanism that responds faster than the provisioning loop. The honest framing is that autoscaling is a cost-optimisation mechanism, not a burst-survival mechanism — it handles the daily ramp and the gradual trend well, and everything faster must be met with capacity that already exists.
Interview question
"Your platform expects unpredictable bursts lasting a few minutes. Walk me through autoscaling, pre-provisioned capacity, serverless, queues and admission control — and tell me which combination you would choose if the bursts are scheduled versus if they are genuinely unpredictable."