Term Kind Topic What it is
Tail Latency p99, p999 metric Performance & Capacity The latency experienced by the slowest small percentage of requests, which is what users and dependent services actually feel.
Tail Latency Amplification Fan-out Tail, Slowest-Component Latency concept Tail Latency A parallel fan-out completes when its slowest component does, so ordinary per-component tails combine into a much worse end-to-end tail.
Tail Latency Amplification concept Tail Latency The effect where a request that fans out to many services experiences the worst case of all of them, making rare slowness common at the user level.
Throughput metric Throughput Work completed per unit time — bounded by the system's narrowest resource, and traded against latency.
Twitter's Timeline Fan-Out case-study Performance & Capacity Twitter precomputes each user's timeline at write time but handles very-high-follower accounts at read time, because neither strategy alone survives both ends of the distribution.
Universal Scalability Law USL, Gunther's Law concept Horizontal vs Vertical Scaling A model showing that throughput rises with concurrency, flattens due to contention, and then falls due to coherency costs.
Utilisation and Queueing Delay concept Queueing Theory The non-linear relationship by which waiting time grows as utilisation approaches one, explaining why systems degrade suddenly rather than gradually.
Utilisation Target practice Queueing Theory The operating point chosen from a latency requirement rather than from cost efficiency, because queueing delay rises non-linearly as utilisation approaches saturation.
WhatsApp's Small-Team Scale case-study Performance & Capacity WhatsApp served hundreds of millions of users with a few dozen engineers by matching one technology choice precisely to the workload and refusing to add anything else.
Working Set Hot Data Set concept Caching for Performance The subset of data actually touched in a given window, whose size against the cache's capacity determines the hit rate and therefore how much load reaches the origin.
Workload Model practice Load Testing A description of the traffic mix, arrival pattern and data distribution a load test reproduces, which determines whether the test's results mean anything.
Write Contention Row Contention, Hot Row concept Performance & Capacity Many concurrent writers competing for the same item of state, which serialises through one lock and cannot be relieved by adding capacity.