Term Kind Topic What it is
Access Pattern Skew Popularity Skew, Zipf Distribution concept Caching for Performance The uneven distribution of requests across keys - which determines whether caching helps at all, and which synthetic load tests systematically fail to reproduce.
Adaptive Admission Level Priority Admission Threshold, Collaborative Load Shedding pattern Queueing Theory A load-dependent priority threshold that a service publishes back to its callers, so requests that would be rejected downstream are never sent and upstream work is not wasted.
Admission Queue Virtual Waiting Room, Login Queue, Front-Door Queue pattern Performance & Capacity Holding excess arrivals in an explicit, communicated queue and admitting them at a rate the system can serve - converting an overload failure into a wait, which is a far better outcome and must be built in advance.
Amdahl's Law concept Profiling & Optimisation The limit on speedup from optimising or parallelising part of a system, set by the proportion of work that remains unchanged.
Bottleneck concept Bottleneck Analysis The single resource that limits system throughput, such that improving anything else produces no gain.
Bottleneck Analysis practice Bottleneck Analysis Finding the single resource that limits the system, because improving anything else changes nothing.
Breaking-Point Testing Test to Failure, Constraint Discovery practice Load Testing Pushing a load test past the target until the system fails, because the purpose of a load test is to locate the constraint and the failure mode rather than to confirm a number.
Cache Hit Ratio Economics concept Caching for Performance The non-linear relationship between cache hit rate and backend load, which determines whether a cache improvement is worth making.
Cache Stampede Thundering Herd, Dog-Piling concept Caching for Performance The surge of identical expensive requests to the origin when a popular cache entry expires and many concurrent callers all miss at once.
Caching Strategy pattern Performance & Capacity The chosen pattern for how a cache is populated, read and invalidated — cache-aside, read-through, write-through or write-behind.
Capacity Lead Time Provisioning Lead Time concept Capacity Modelling The delay between deciding capacity is needed and having it available, which determines whether a capacity model can forecast at all or must instead buy optionality.
Capacity Modelling practice Capacity Modelling Predicting the resources a workload will need, from measured unit costs and a demand forecast, with deliberate headroom.
Collective Communication Cost Allreduce Overhead, Synchronisation Cost at Scale concept Throughput The synchronisation term that determines how far a distributed training or inference job can scale, made of a bandwidth component that is nearly flat in participant count and a latency component that is not.
Concurrency concept Concurrency How many operations are in flight at once — the quantity that actually saturates a system, and the one most worth limiting explicitly.
Connection Pool concept Performance & Capacity A fixed set of reusable database connections shared by an application's requests, and one of the most common hidden capacity ceilings.
Connection Pool Sizing practice Connection Pooling Choosing how many concurrent connections a service holds to a datastore, where both too many and too few cause outages.
Coordinated Omission Closed-Loop Measurement Bias, Omitted Latency Samples concept Load Testing The measurement error in which a load generator stops issuing requests while the system under test is stalled, so the slow samples that a real arrival process would have produced are never recorded.
Database Performance practice Database Performance The interventions that resolve database bottlenecks, in the order they should be attempted.
Deadline-Derived Queue Depth Deadline-Bounded Queue concept Concurrency The rule that a request queue's maximum depth comes from the client's deadline and the service's completion rate, so the queue can never hold work that will expire before it is served.
Demand Forecasting practice Capacity Modelling Projecting future load from historical trends, business plans and known events, and translating it into resource requirements with explicit lead times.
Efficient Concurrency Point Optimal Pool Size, Contention Threshold concept Connection Pooling The concurrency level beyond which adding parallel work reduces total throughput, because contention costs more than the parallelism gains - and the reason a larger connection pool often makes latency worse.
Ephemeral Port Exhaustion Source Port Starvation, TIME_WAIT Exhaustion concept Network Performance Tuning The failure in which a host runs out of source ports for new connections to one destination, producing connection errors that look like the remote being down while the remote is healthy.
Fan-Out Latency Amplification concept Tail Latency The effect by which a request that depends on many parallel sub-requests is governed by the slowest of them, so rare slowness becomes common at the user level.
Flame Graph tool Profiling & Optimisation A visualisation of sampled stack traces in which width represents time spent, used to identify where a program's execution actually goes.
Goodput Useful Throughput metric Throughput The rate of completed work that someone is still waiting for, which diverges from throughput under overload and is the only rate that tracks what users receive.
Google Maps and Planetary-Scale Spatial Serving case-study Performance & Capacity Map serving is fast because almost nothing is computed on request — the world is precomputed into a pyramid of tiles, and space is indexed onto a one-dimensional curve.
Horizontal vs Vertical Scaling Scale Out vs Scale Up concept Performance & Capacity Adding more machines versus making one machine bigger — and the fact that vertical is underrated for stateful tiers.
In-Memory Reservation Service Inventory Reservation Service, Single-Owner Counter, Memory-Speed Allocation pattern Database Performance Moving a heavily-contended counter or inventory pool into a single-owner in-memory service that serialises decisions at memory speed and persists asynchronously, trading a small durability window for orders of…
Index Write Amplification Index Maintenance Cost, Secondary Index Overhead concept Database Performance The multiplication of write work caused by secondary indexes, where every insert and every update to an indexed column must also maintain each index tree inside the same transaction.
Latency metric Latency How long one operation takes — a distribution, never a number, and dominated by its tail in any system with fan-out.
Latency Budget Decomposition practice Latency Allocating a total response-time target across the components of a request path, so each layer has an explicit share and overruns are attributable.
Little's Law concept Performance & Capacity In a stable system, the average number of items in it equals the arrival rate times the average time each spends in it — L = λW.
Little's Law Applied to Pools practice Connection Pooling Using L = λW to size connection and thread pools from measured throughput and latency rather than from a default.
Little's Law in Practice metric Little's Law L = λW — concurrency equals arrival rate times latency — and the reason a slowdown becomes an outage.
Load Testing practice Performance & Capacity Driving a system with realistic traffic at a target volume to verify it meets its performance targets before real users do.
Load Testing in Practice practice Load Testing Verifying behaviour under expected load — where realism of workload and data matters more than the number of virtual users.
Network Performance Tuning practice Network Performance Tuning The application-level and connection-level changes that actually improve networked performance, in order of effect.
Overload Recovery Time Overload Settling Time metric Stress Testing Wall-clock time from the moment load returns to normal until the system is fully healthy again, which is set by spare capacity rather than by service speed.
Parallelism and Concurrency concept Concurrency The distinction between structuring work so it can make progress independently and actually executing work simultaneously on multiple cores.
Payload Compression Trade-off concept Network Performance Tuning The exchange of CPU time for reduced transfer time, which is favourable on constrained networks and harmful on fast local ones.
Peak Event Readiness practice Peak Event Readiness Preparing for a known, dated, high-stakes traffic event — where the discipline is as much organisational as technical.
Peak Readiness Rehearsal Pre-Event Verification, Scheduled Peak Drill practice Peak Event Readiness The set of pre-peak actions and verifications - warm caches, warm pools, confirmed dependency headroom, a change freeze, a rehearsed shedding path - without which correct capacity still fails.
Percentile Latency metric Latency Latency expressed as the value below which a given proportion of requests fall, used because averages conceal the behaviour that users notice.
Performance Budget practice Performance Budgets A stated numeric limit on a performance characteristic, enforced automatically so that regressions fail a build rather than accumulating.
Performance Budget for Backends practice Performance Budgets A latency or resource allowance allocated per component of a request path, so that the end-to-end target is composed rather than hoped for.
Pre-Scaling Scheduled Scaling, Warm Provisioning, Capacity Staging practice Peak Event Readiness Provisioning capacity in advance of a known event rather than relying on autoscaling, because the arrival spike at a scheduled moment is faster than any reactive control loop can respond to.
Profiling and Optimisation practice Profiling & Optimisation The discipline of measuring before changing, optimising the dominant term, and stopping when the objective is met.
Query Plan Regression concept Database Performance A sudden latency increase caused by the database optimiser choosing a different execution plan for an unchanged query, typically after statistics or data volume change.
Queue Length Estimation practice Little's Law Using the relationship between arrival rate, residence time and items in the system to size pools, predict backlogs and sanity-check capacity claims.
Queueing Delay concept Queueing Theory The time a request spends waiting rather than being served, which rises non-linearly as utilisation approaches capacity.
Resource Saturation Audit practice Bottleneck Analysis A systematic checklist for locating resource bottlenecks by examining utilisation, saturation and errors for every resource in the system.
Retail Peak: Designing for Black Friday Black Friday Readiness case-study Peak Event Readiness A retailer's annual peak can be an order of magnitude above normal, arrives in minutes, and cannot be rescheduled — which makes it a distinct engineering discipline.
Saturation Point concept Throughput The load level beyond which additional demand produces queueing and latency growth rather than additional completed work.
Scale Cube concept Horizontal vs Vertical Scaling A model describing three independent axes of scaling — cloning, functional decomposition, and data partitioning — each addressing a different limit.
Segment-Isolated Serving Outlier Tenant Path, Heavy-User Isolation pattern Tail Latency Giving a consistently slow population its own path - pool, precomputed view, or query strategy - rather than optimising the shared path to accommodate it.
Soak Test practice Soak Testing Running sustained realistic load for hours or days to expose defects that accumulate over time rather than appearing under peak load.
Soak Testing practice Soak Testing Running at sustained realistic load for hours or days to find the failures that only appear with time.
Stress Test practice Stress Testing Driving load beyond expected capacity to observe how the system behaves at and past its breaking point.
Stress Testing practice Stress Testing Deliberately exceeding capacity to find where the system breaks and — more importantly — how.
Tail Composition Effect Fan-Out Tail Amplification, Page-Level Tail concept Tail Latency The multiplication that turns a small per-request slow rate into a large user-visible slow rate, because a single user action makes many requests and is only as fast as its slowest one.