Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
3 to work through
-
intermediate Multiple choice
A colleague spent a week optimising a function and the endpoint is 2% faster. What went wrong in the approach?
2 min answer -
intermediate
A team has profiled a service and found the top three functions by CPU time. Why is optimising them often the wrong next step?
2 min answer -
advanced
Discord's Read States service, written in Go, showed latency spikes every two minutes like clockwork. The team had written it carefully with very few allocations, and the spikes appeared regardless of load. Their published account from 2020 explains the cause and the rewrite that followed. What was happening, and what does it teach about periodic latency?
3 min answer
3 terms in this topic
Amdahl's Law
The limit on speedup from optimising or parallelising part of a system, set by the proportion of work that remains unchanged.
toolFlame Graph
A visualisation of sampled stack traces in which width represents time spent, used to identify where a program's execution actually goes.
practiceProfiling and Optimisation
The discipline of measuring before changing, optimising the dominant term, and stopping when the objective is met.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.