Performance & Capacity
General material on performance and capacity engineering.
6 to work through
-
intermediate
A search engine holds its index entirely in memory to guarantee predictable latency. What does that buy, what does it cost, and when is it the wrong choice?
2 min answer -
intermediate
Peak trading day is six weeks away and expected to be four times normal traffic. What do you do in those six weeks?
3 min answer -
intermediate
Precompute every user's timeline at write time, or assemble it at read time? Explain why the answer for a social feed is neither.
2 min answer -
intermediate
Your system handles 1,000 requests per second today. Marketing says a campaign will bring 10,000 next month. What breaks first, and how do you find out?
2 min answer -
advanced
A distributed training job runs at 90% scaling efficiency on 512 GPUs. At 1,024 GPUs it drops to 60%. Walk me through where the time went, what you would measure, and what you would try.
3 min answer -
advanced
Millions of customers attempt to buy a small number of items at a scheduled instant. Which performance constraint dominates, and why do conventional scaling techniques fail against it?
2 min answer
11 terms in this topic
Admission Queue
Holding excess arrivals in an explicit, communicated queue and admitting them at a rate the system can serve - converting an overload failure into a …
patternCaching Strategy
The chosen pattern for how a cache is populated, read and invalidated — cache-aside, read-through, write-through or write-behind.
conceptConnection Pool
A fixed set of reusable database connections shared by an application's requests, and one of the most common hidden capacity ceilings.
case-studyGoogle Maps and Planetary-Scale Spatial Serving
Map serving is fast because almost nothing is computed on request — the world is precomputed into a pyramid of tiles, and space is indexed onto a one…
conceptHorizontal vs Vertical Scaling
Adding more machines versus making one machine bigger — and the fact that vertical is underrated for stateful tiers.
conceptLittle's Law
In a stable system, the average number of items in it equals the arrival rate times the average time each spends in it — L = λW.
practiceLoad Testing
Driving a system with realistic traffic at a target volume to verify it meets its performance targets before real users do.
metricTail Latency
The latency experienced by the slowest small percentage of requests, which is what users and dependent services actually feel.
case-studyTwitter's Timeline Fan-Out
Twitter precomputes each user's timeline at write time but handles very-high-follower accounts at read time, because neither strategy alone survives …
case-studyWhatsApp's Small-Team Scale
WhatsApp served hundreds of millions of users with a few dozen engineers by matching one technology choice precisely to the workload and refusing to …
conceptWrite Contention
Many concurrent writers competing for the same item of state, which serialises through one lock and cannot be relieved by adding capacity.
Neighbouring topics
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.