Cloud Architecture
General material on designing for cloud platforms.
8 to work through
-
intermediate
A cost-sensitive education platform must serve large enrolment spikes and heavy video delivery on a tight budget. Which architectural choices give the largest cost reduction per unit of effort?
2 min answer -
intermediate
A team tells you backups run nightly and are retained for 30 days. What is missing from that answer?
2 min answer -
intermediate Multiple choice
A team wants to build a new internal API on serverless functions. It will serve steady traffic of about 200 requests per second during business hours. What do you advise?
2 min answer -
advanced Multiple choice
A deployment platform's control plane goes down. What should happen to customer applications already serving traffic, and what does that requirement force into the architecture?
2 min answer -
advanced Multiple choice
A real-time conferencing platform must serve unpredictable demand that once grew thirtyfold in ten weeks, where media relaying is bandwidth-intensive and latency-critical. Should it run on public cloud, its own data centres, or both - and what decides?
2 min answer -
advanced
In the 2017 AWS S3 outage, the status page could not report the outage because it depended on S3. What does that tell you about designing status and control systems?
2 min answer -
advanced
The December 2021 us-east-1 event began with internal network congestion that impaired the region's control plane, and customers found they could not even read the service health dashboard. What are the three durable architectural lessons, independent of the specific bug?
2 min answer -
advanced
The business asks for "multi-region" after a regional outage. Before agreeing, what do you need to establish, and what are you actually signing up for?
2 min answer
19 terms in this topic
Autoscaling
Adding and removing capacity automatically in response to a demand signal, to track load without paying for peak all the time.
conceptAvailability Zone
One or more physically separate data centres inside a cloud region, with independent power, cooling and network, connected by low-latency links.
practiceBackup Strategy
A plan for what is copied, how often, where to, how long it is kept, and — the part that decides whether it is real — how the restore is verified.
conceptBlast Radius
The set of things that break, or become reachable, when one component fails or is compromised.
practiceBlast Radius Reduction
The set of deliberate partitions — accounts, regions, zones, cells, tenants, deployment stages — that bound how far any single failure or compromise …
conceptBurst Capacity
Capacity acquired quickly for short-lived demand above the baseline, deliberately priced higher than owned capacity in exchange for immediate availability.
conceptControl Plane and Data Plane
The separation between the machinery that makes changes to a system and the machinery that serves its traffic.
conceptControl Plane vs Data Plane
The separation between the system that manages configuration and the system that serves requests, which have very different reliability requirements.
practiceEgress-First Cost Modelling
Establishing which cost category actually dominates before optimising anything - because in media-heavy consumer platforms the order is usually bandw…
practiceInfrastructure as Code
Defining infrastructure in version-controlled declarative files that a tool reconciles against the real environment.
toolKubernetes
A container orchestrator that continuously reconciles the running state of a cluster towards a declared desired state.
conceptManaged Service
A capability the provider operates — provisioning, patching, backup, scaling and failover — leaving you the configuration and the data.
conceptMulti-Cloud
Deliberately running across more than one cloud provider — a decision with a much higher cost than the lock-in it is usually adopted to avoid.
conceptObject Storage
Flat, HTTP-addressable storage for immutable blobs with rich metadata — effectively unlimited, cheap, and not a filesystem.
conceptServerless
A model where the provider allocates and scales compute per request, and you are billed for execution rather than for provisioned capacity.
case-studySlack's Cellular Migration
After repeated availability-zone-level incidents, Slack rebuilt its infrastructure into per-zone cells with the ability to drain traffic away from a …
conceptStatic Stability
Designing a system so it continues operating correctly on its existing configuration during a failure, requiring no control-plane action to survive -…
practiceWell-Architected Review
A structured self-assessment of a workload against defined pillars — operational excellence, security, reliability, performance, cost, and sustainability.
case-studyZoom's Pandemic Scale-Up
Zoom grew from around 10 million to over 300 million daily meeting participants in roughly three months, absorbed by a hybrid architecture and a dist…
Neighbouring topics
Compute Models
Instances, containers and functions, and what each is priced and shaped for.
Cloud Storage
Object, block and file storage, and the access patterns each suits.
Cloud Databases
Managed relational, key-value, document and analytical services.
Containers
Images, registries, immutability and the deployment model they enable.
Kubernetes
The reconciliation loop, and whether the workload needs what it provides.
Serverless
Scale to zero, per-request billing, cold starts and connection limits.
Autoscaling
Signals, delays and bounds — and the maximum that caps a runaway bill.
Load Balancing
Distributing traffic, health checking, and removing failures from rotation.
Multi-Region Architecture
Surviving a region, and the data consistency price of doing so.
Availability Zones
The unit of correlated physical failure, and what zones do not protect against.
Disaster Recovery
Backup-restore, pilot light, warm standby and active-active postures.
Backup Strategies
Scope, immutability, separation, and the restore drill that makes it real.
Infrastructure as Code
Declarative infrastructure, drift, state files and rebuild-from-empty.
Landing Zones
A governed foundation of accounts, network, identity and guardrails.
Managed Services
Which operational responsibilities actually transfer, and which do not.
Cloud Migration
Per-application disposition, sequencing and the capability change underneath.
Multi-Cloud
Best-of-breed, portfolio and portable — three very different costs.
Edge Computing
Moving compute towards the user, and what cannot follow it.
Cloud Governance
Preventive policy, tagging, quotas and cost and security guardrails.