Autoscaling
Signals, delays and bounds — and the maximum that caps a runaway bill.
5 to work through
-
advanced
A live-streaming platform expects viewer count to increase tenfold within minutes when a major event begins. Its services autoscale on CPU. Predict what happens and what the design should be instead.
2 min answer -
advanced
A service autoscales on CPU. During incidents it never scales out, even as latency triples. Why, and what would you scale on instead?
2 min answer -
advanced
A traffic ramp after a quiet period saturates a managed network component that scales automatically but not fast enough. Degradation cascades. Analyse.
2 min answer -
advanced
A workload has extreme bursts lasting a few minutes, triggered by scheduled real-world events. Should the architecture use autoscaling, pre-provisioned capacity, queues, caching, admission control, serverless, or a combination?
2 min answer -
advanced
Your capacity plan relies on autoscaling to handle traffic spikes. Why might that fail exactly when needed?
2 min answer
5 terms in this topic
Pre-Scaling
Provisioning capacity before demand arrives, based on a forecast or a known event, because autoscaling cannot respond faster than its provisioning loop.
practiceScheduled Capacity
Provisioning ahead of a known event rather than reacting to its traffic - the correct primary control for any spike whose rise time is shorter than a…
case-studySlack 2021: When Autoscaling Cannot Keep Up
The first Monday back after the holidays produced a traffic ramp that outpaced the scaling behaviour of a managed network component, and the degradat…
patternTarget Tracking Scaling
An autoscaling policy that adds or removes capacity to hold a chosen metric near a target value, like a thermostat, rather than reacting to threshold…
patternWarm Pool
Pre-initialised instances held in a stopped or standby state so that scaling out skips boot and application warm-up.
Neighbouring topics
Cloud Architecture
General material on designing for cloud platforms.
Compute Models
Instances, containers and functions, and what each is priced and shaped for.
Cloud Storage
Object, block and file storage, and the access patterns each suits.
Cloud Databases
Managed relational, key-value, document and analytical services.
Containers
Images, registries, immutability and the deployment model they enable.
Kubernetes
The reconciliation loop, and whether the workload needs what it provides.
Serverless
Scale to zero, per-request billing, cold starts and connection limits.
Load Balancing
Distributing traffic, health checking, and removing failures from rotation.
Multi-Region Architecture
Surviving a region, and the data consistency price of doing so.
Availability Zones
The unit of correlated physical failure, and what zones do not protect against.
Disaster Recovery
Backup-restore, pilot light, warm standby and active-active postures.
Backup Strategies
Scope, immutability, separation, and the restore drill that makes it real.
Infrastructure as Code
Declarative infrastructure, drift, state files and rebuild-from-empty.
Landing Zones
A governed foundation of accounts, network, identity and guardrails.
Managed Services
Which operational responsibilities actually transfer, and which do not.
Cloud Migration
Per-application disposition, sequencing and the capability change underneath.
Multi-Cloud
Best-of-breed, portfolio and portable — three very different costs.
Edge Computing
Moving compute towards the user, and what cannot follow it.
Cloud Governance
Preventive policy, tagging, quotas and cost and security guardrails.