Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
4 to work through
-
advanced
A fashion marketplace's average latency is excellent but p99 is unacceptable for a small, commercially important customer segment. How should tracing, segmentation, queue analysis and dependency timing guide the investigation?
2 min answer -
advanced
A request fans out to 100 services in parallel. Each responds within 10 ms for 99% of calls. What fraction of user requests are slow, and what do you do?
2 min answer -
advanced
A video platform's request path fans out to twenty backend services in parallel. Each has a p99 of 50 ms. Why is the overall p99 far worse than 50 ms, and what fixes it?
2 min answer -
advanced
You add hedged requests to cut tail latency. It works well, then during a traffic peak the service collapses. Explain.
2 min answer
5 terms in this topic
Fan-Out Latency Amplification
The effect by which a request that depends on many parallel sub-requests is governed by the slowest of them, so rare slowness becomes common at the u…
patternSegment-Isolated Serving
Giving a consistently slow population its own path - pool, precomputed view, or query strategy - rather than optimising the shared path to accommodate it.
conceptTail Composition Effect
The multiplication that turns a small per-request slow rate into a large user-visible slow rate, because a single user action makes many requests and…
conceptTail Latency Amplification
A parallel fan-out completes when its slowest component does, so ordinary per-component tails combine into a much worse end-to-end tail.
conceptTail Latency Amplification
The effect where a request that fans out to many services experiences the worst case of all of them, making rare slowness common at the user level.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.