Latency
Distributions rather than averages, and the floors physics imposes.
3 to work through
-
advanced
A collaborative editor must feel instantaneous to users on different continents, where physics imposes a floor of roughly 150 ms round trip. How is the latency requirement met at all?
2 min answer -
advanced
A load generator reports a p99 of 30 ms at 5,000 rps. In production at the same rate, users report multi-second stalls, and the service's own latency metrics agree with the load test. Both measurements are honest. What is being missed?
3 min answer -
advanced
A page makes 20 parallel backend calls, each with a p99 of 1 second and a p50 of 40 ms. What is the page's latency profile?
2 min answer
3 terms in this topic
Latency
How long one operation takes — a distribution, never a number, and dominated by its tail in any system with fan-out.
practiceLatency Budget Decomposition
Allocating a total response-time target across the components of a request path, so each layer has an explicit share and overruns are attributable.
metricPercentile Latency
Latency expressed as the value below which a given proportion of requests fall, used because averages conceal the behaviour that users notice.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.