Term Kind Topic What it is
Point-in-Time Recovery PITR concept Backup Strategies Restoring a database to any moment within a retention window by replaying transaction logs onto a base backup, rather than only to a snapshot boundary.
Pre-Scaling Scheduled Scaling, Predictive Scaling practice Autoscaling Provisioning capacity before demand arrives, based on a forecast or a known event, because autoscaling cannot respond faster than its provisioning loop.
Preventive Guardrail Policy as Prevention, Deny-by-Default Control practice Cloud Governance A control that makes a non-compliant action impossible, rather than detecting it afterwards and generating a ticket.
Provisioned vs Serverless Capacity concept Cloud Databases Paying for a fixed database size continuously, versus paying for capacity consumed with automatic scaling — a crossover decision driven by duty cycle.
Regional Evacuation Region Failover Drill practice Multi-Region Architecture The deliberate, rehearsed act of shifting all traffic out of a region — treated as a routine operation rather than an emergency procedure.
Resource Requests and Limits concept Kubernetes The declared minimum a container is guaranteed (request) and the maximum it may consume (limit) — the two numbers that determine scheduling, packing and throttling.
Restore Drill practice Backup Strategies A scheduled, timed exercise of restoring from backup into a clean environment — the only thing that converts a backup from a hope into a control.
Revocation Path Fast-Propagation Channel, Fail-Closed Config pattern Cloud Governance A separate, short-TTL, fail-closed channel for security-critical configuration changes, coexisting with an eventually-consistent cached channel for ordinary configuration.
Scheduled Capacity Predictive Provisioning, Calendar-Driven Scaling practice Autoscaling Provisioning ahead of a known event rather than reacting to its traffic - the correct primary control for any spike whose rise time is shorter than an autoscaler's reaction time.
Serverless concept Cloud Architecture A model where the provider allocates and scales compute per request, and you are billed for execution rather than for provisioned capacity.
Serverless Cold Start concept Serverless The additional latency when a function invocation must allocate and initialise a new execution environment rather than reusing a warm one.
Serverless Compute in Practice Functions as a Service, FaaS concept Serverless Compute billed per invocation with no capacity to manage — excellent for bursty low-duty-cycle work, expensive and constrained for sustained load.
Service Control Policy SCP, Azure Policy, Organization Policy tool Landing Zones An organisation-level guardrail that limits what any identity in an account may do, regardless of the permissions granted within that account.
Service Quota Service Limit concept Cloud Governance A per-account, per-region cap on how much of a resource may be used — a common and easily-avoided cause of scaling failures and DR failures.
Shared Responsibility Model concept Managed Services The division of duties between provider and customer, which shifts with the service model and is routinely misunderstood in the customer's disfavour.
Slack 2021: When Autoscaling Cannot Keep Up Slack January 2021 Outage case-study Autoscaling The first Monday back after the holidays produced a traffic ramp that outpaced the scaling behaviour of a managed network component, and the degradation cascaded.
Slack's Cellular Migration case-study Cloud Architecture After repeated availability-zone-level incidents, Slack rebuilt its infrastructure into per-zone cells with the ability to drain traffic away from a failing zone in minutes.
Static Stability Fail Static, Configuration Inertia, Control-Plane Independence concept Cloud Architecture Designing a system so it continues operating correctly on its existing configuration during a failure, requiring no control-plane action to survive - because control planes are the part most likely to be unava…
Storage Tiering Lifecycle Management, Hot-Warm-Cold Storage practice Cloud Storage Placing objects in storage classes according to access probability, so cost tracks usage rather than volume.
Superseded-Work Cancellation Build Cancellation, Queue Deduplication practice Containers Cancelling queued work that a later submission has made irrelevant - usually the largest single capacity saving available during a burst, and almost free to implement.
Tagging Strategy practice Cloud Governance A defined, enforced set of metadata labels applied to every resource, without which cost allocation, ownership and lifecycle automation are all impossible.
Target Tracking Scaling pattern Autoscaling An autoscaling policy that adds or removes capacity to hold a chosen metric near a target value, like a thermostat, rather than reacting to threshold breaches.
Terraform State IaC State File concept Infrastructure as Code The file mapping declared resources to real infrastructure, without which the tool cannot tell what it already created — and which becomes critical infrastructure in its own right.
Verified Recovery Restore Assurance, Tested Recoverability metric Backup Strategies The measure that matters in data protection - the proportion of backups demonstrably restorable within the stated RTO - as opposed to backup success rate, which is near 100% even where recovery is impossible.
Warm Pool pattern Autoscaling Pre-initialised instances held in a stopped or standby state so that scaling out skips boot and application warm-up.
Warm Standby practice Disaster Recovery A continuously-replicated, scaled-down running copy of a system in a second location, ready to take traffic within minutes.
Well-Architected Review practice Cloud Architecture A structured self-assessment of a workload against defined pillars — operational excellence, security, reliability, performance, cost, and sustainability.
Zonal vs Regional Services concept Availability Zones Whether a cloud resource lives in one availability zone or is inherently spread across several — a property that determines what a zone failure takes with it.
Zoom's Pandemic Scale-Up case-study Cloud Architecture Zoom grew from around 10 million to over 300 million daily meeting participants in roughly three months, absorbed by a hybrid architecture and a distributed media routing design.