Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
5 to work through
-
intermediate
A service is slow and the team disagrees about the cause. What method identifies the actual bottleneck rather than the suspected one?
2 min answer -
advanced
A marketplace's search page is slow. CPU is moderate, memory is fine, and the database reports healthy query times. How do you find the bottleneck?
2 min answer -
advanced
A service's p99 latency has tripled over three months with no single obvious change. How do you investigate?
2 min answer -
advanced
CPU, memory, disk and network all look healthy, and the service will not go faster. What are you looking for?
2 min answer -
advanced
p99 latency on the booking endpoint jumped from 300 ms to 4 s at 09:00 today. Error rate is normal. Walk through your diagnosis.
2 min answer
3 terms in this topic
Bottleneck
The single resource that limits system throughput, such that improving anything else produces no gain.
practiceBottleneck Analysis
Finding the single resource that limits the system, because improving anything else changes nothing.
practiceResource Saturation Audit
A systematic checklist for locating resource bottlenecks by examining utilisation, saturation and errors for every resource in the system.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.