A service's p99 is 300 ms and an hour-long sampling CPU profile shows no single function above 4% of samples. The flame graph is wide and shallow everywhere. What does that tell you to do next?
Show the full answer Hide the answer
The diagnosis
A flat profile is a result, not a failed measurement. It means the cost is spread across the per-request path rather than concentrated in an algorithm: framework middleware, serialisation and deserialisation, logging, allocation and garbage collection, context propagation, ORM materialisation, TLS. Each costs 1–4% and the sum is the bill.
Amdahl's law then tells you what local optimisation can achieve. The best possible outcome from perfecting a 4% frame is a 4% improvement, and a realistic outcome is 1–2%. Three such projects cost a quarter and deliver a change no user can perceive. The lever that works on a flat profile is fewer units of work: fewer requests (batch the client's calls, cache whole responses), or less work per request (drop a serialisation hop, remove a middleware layer, stop logging the payload).
The step before that, which the question is baiting
Reconcile on-CPU time per request against wall-clock latency before concluding anything from the shape of the flame graph. Divide total CPU seconds by requests served: if the service burns 12 ms of CPU per request while p99 is 300 ms, then 96% of the latency is off-CPU — blocked on a lock, on IO, on a connection-pool checkout, on garbage-collection pauses or in the scheduler's run queue. The profile is flat partly because there is barely any CPU work to concentrate.
That changes the whole investigation: off-CPU analysis (scheduler and lock profiling, wait-time spans in traces) rather than CPU optimisation. The misleading signal is the flame graph itself, which is a picture of where CPU went and silently implies that CPU is where time went.
The fix, in order
- Account for the gap between CPU milliseconds per request and wall-clock latency.
- If on-CPU time dominates and the profile is flat, attack request count and per-request overhead.
- If off-CPU time dominates, profile off-CPU and expect a queue, a lock or a pool.
The alert that would have caught it earlier
CPU milliseconds per request, charted beside latency. A rising latency with flat CPU per request is a queueing or blocking problem; rising CPU per request is a code or payload problem. One ratio distinguishes two investigations that look identical on a latency dashboard.
When this is the wrong conclusion
A profile can also be flat because the sampler cannot see the stacks. Missing frame pointers, aggressive inlining, a JIT without a symbol map or a native library without debug information all produce broad, shallow graphs that are measurement artefacts. Check that symbolisation is complete before believing the shape. And if CPU per request genuinely accounts for the latency and the profile is genuinely flat, the honest answer may be that the code is already efficient and the system simply needs more instances, which is a capacity decision rather than an optimisation one.
Why the other options fail
- Optimise the top frame. At 4% of CPU, and with CPU only 12 ms of a 300 ms request, the ceiling is roughly 0.5 ms, or 0.16% of p99. This is the right move when a profile has a 40% frame, which is exactly why people reach for it when it has none.
- Raise the sampling rate. More samples refine the estimates; they cannot create concentration that is not there. An hour at 99 Hz is already hundreds of thousands of samples. It is the right move for a noisy profile from a 30-second window, not for a flat one from an hour.
- Switch to instrumentation. Per-call instrumentation adds overhead to the thing being measured and still returns a flat distribution, just more precisely. It earns its cost when you need exact call counts, such as proving an N+1 query, not when you need to know where wall-clock time went.