Term Kind Topic What it is
Access Pattern Skew Popularity Skew, Zipf Distribution concept Caching for Performance The uneven distribution of requests across keys - which determines whether caching helps at all, and which synthetic load tests systematically fail to reproduce.
Amdahl's Law concept Profiling & Optimisation The limit on speedup from optimising or parallelising part of a system, set by the proportion of work that remains unchanged.
Bottleneck concept Bottleneck Analysis The single resource that limits system throughput, such that improving anything else produces no gain.
Cache Hit Ratio Economics concept Caching for Performance The non-linear relationship between cache hit rate and backend load, which determines whether a cache improvement is worth making.
Cache Stampede Thundering Herd, Dog-Piling concept Caching for Performance The surge of identical expensive requests to the origin when a popular cache entry expires and many concurrent callers all miss at once.
Capacity Lead Time Provisioning Lead Time concept Capacity Modelling The delay between deciding capacity is needed and having it available, which determines whether a capacity model can forecast at all or must instead buy optionality.
Collective Communication Cost Allreduce Overhead, Synchronisation Cost at Scale concept Throughput The synchronisation term that determines how far a distributed training or inference job can scale, made of a bandwidth component that is nearly flat in participant count and a latency component that is not.
Concurrency concept Concurrency How many operations are in flight at once — the quantity that actually saturates a system, and the one most worth limiting explicitly.
Connection Pool concept Performance & Capacity A fixed set of reusable database connections shared by an application's requests, and one of the most common hidden capacity ceilings.
Coordinated Omission Closed-Loop Measurement Bias, Omitted Latency Samples concept Load Testing The measurement error in which a load generator stops issuing requests while the system under test is stalled, so the slow samples that a real arrival process would have produced are never recorded.
Deadline-Derived Queue Depth Deadline-Bounded Queue concept Concurrency The rule that a request queue's maximum depth comes from the client's deadline and the service's completion rate, so the queue can never hold work that will expire before it is served.
Efficient Concurrency Point Optimal Pool Size, Contention Threshold concept Connection Pooling The concurrency level beyond which adding parallel work reduces total throughput, because contention costs more than the parallelism gains - and the reason a larger connection pool often makes latency worse.
Ephemeral Port Exhaustion Source Port Starvation, TIME_WAIT Exhaustion concept Network Performance Tuning The failure in which a host runs out of source ports for new connections to one destination, producing connection errors that look like the remote being down while the remote is healthy.
Fan-Out Latency Amplification concept Tail Latency The effect by which a request that depends on many parallel sub-requests is governed by the slowest of them, so rare slowness becomes common at the user level.
Horizontal vs Vertical Scaling Scale Out vs Scale Up concept Performance & Capacity Adding more machines versus making one machine bigger — and the fact that vertical is underrated for stateful tiers.
Index Write Amplification Index Maintenance Cost, Secondary Index Overhead concept Database Performance The multiplication of write work caused by secondary indexes, where every insert and every update to an indexed column must also maintain each index tree inside the same transaction.
Little's Law concept Performance & Capacity In a stable system, the average number of items in it equals the arrival rate times the average time each spends in it — L = λW.
Parallelism and Concurrency concept Concurrency The distinction between structuring work so it can make progress independently and actually executing work simultaneously on multiple cores.
Payload Compression Trade-off concept Network Performance Tuning The exchange of CPU time for reduced transfer time, which is favourable on constrained networks and harmful on fast local ones.
Query Plan Regression concept Database Performance A sudden latency increase caused by the database optimiser choosing a different execution plan for an unchanged query, typically after statistics or data volume change.
Queueing Delay concept Queueing Theory The time a request spends waiting rather than being served, which rises non-linearly as utilisation approaches capacity.
Saturation Point concept Throughput The load level beyond which additional demand produces queueing and latency growth rather than additional completed work.
Scale Cube concept Horizontal vs Vertical Scaling A model describing three independent axes of scaling — cloning, functional decomposition, and data partitioning — each addressing a different limit.
Tail Composition Effect Fan-Out Tail Amplification, Page-Level Tail concept Tail Latency The multiplication that turns a small per-request slow rate into a large user-visible slow rate, because a single user action makes many requests and is only as fast as its slowest one.
Tail Latency Amplification Fan-out Tail, Slowest-Component Latency concept Tail Latency A parallel fan-out completes when its slowest component does, so ordinary per-component tails combine into a much worse end-to-end tail.
Tail Latency Amplification concept Tail Latency The effect where a request that fans out to many services experiences the worst case of all of them, making rare slowness common at the user level.
Universal Scalability Law USL, Gunther's Law concept Horizontal vs Vertical Scaling A model showing that throughput rises with concurrency, flattens due to contention, and then falls due to coherency costs.
Utilisation and Queueing Delay concept Queueing Theory The non-linear relationship by which waiting time grows as utilisation approaches one, explaining why systems degrade suddenly rather than gradually.
Working Set Hot Data Set concept Caching for Performance The subset of data actually touched in a given window, whose size against the cache's capacity determines the hit rate and therefore how much load reaches the origin.
Write Contention Row Contention, Hot Row concept Performance & Capacity Many concurrent writers competing for the same item of state, which serialises through one lock and cannot be relieved by adding capacity.