Flat Profile
also called Diffuse Cost Profile
A sampled profile with no dominant frame, which says that cost is spread across per-request overhead or spent off CPU entirely - so local optimisation cannot help and the unit of work itself must change.
An engineer profiles a slow service for an hour and finds nothing above 4% of samples. The flame graph is wide and shallow, and the usual reaction is that the sampling rate must be too low.
A flat profile is a result, not a failed measurement. It says the cost is distributed: middleware, serialisation, logging, allocation and garbage collection, context propagation, ORM materialisation, TLS. Each costs 1–4%; their sum is the bill. Or it says something more important — that the latency is not on CPU at all, and the graph is flat because there is barely any CPU work to concentrate.
Why it matters
Amdahl's law (1967) sets the ceiling. Perfecting a 4% frame yields at most 4%, and realistically 1–2%, so three such projects consume a quarter and produce a change no user can perceive. Teams start them anyway, because the top of a profile is where optimisation starts and nobody checks whether the top is tall.
The second reason matters more. Dividing total CPU seconds by requests served gives CPU milliseconds per request. If a service burns 12 ms of CPU inside a 300 ms p99, then 96% of the latency is off CPU — blocked on a lock, on IO, on a pool checkout, on GC pauses, or waiting in the run queue. The flame graph is the misleading signal: it shows where CPU went and implies that CPU is where time went.
Implementation patterns
- Reconcile CPU per request against wall-clock latency before interpreting any profile. One division decides which of two different investigations to run.
- If on-CPU time dominates and the profile is flat, reduce units of work: batch the client's calls, cache whole responses, remove a middleware layer, stop serialising unread fields.
- If off-CPU time dominates, profile off CPU — scheduler and lock profiling, blocked-time stacks, wait spans in traces — and the output names a queue, a lock or a pool.
- Chart CPU milliseconds per request beside latency. Rising latency with flat CPU per request is a queueing problem; rising CPU per request is a code or payload problem.
- Verify symbolisation, and attribute samples per endpoint. Missing frame pointers, inlining or a stripped library produce artificial flatness, and a flat process-wide profile can hide one concentrated endpoint.
Industry example
The shape is routine in mature service estates where a request passes through authentication, a tracing interceptor, validation, deserialisation, an ORM, metrics emission, serialisation and TLS. Each layer was individually defensible, and together they are the latency. The industry's responses are all unit-of-work reductions rather than local optimisations: thinner serialisation formats, batch and streaming APIs that amortise the per-request path, aggregation layers that cut request count, and sampled telemetry. Off-CPU analysis became an established method precisely because CPU flame graphs answer the wrong question for IO-bound services.
Failure scenarios
- A quarter spent optimising a 4% frame, delivering 1% and burning the team's credibility for the next performance proposal.
- A conclusion drawn from a broken profile, where the flatness was missing symbols all along.
- Off-CPU time misread as efficiency: "we only use 12 ms of CPU per request" as good news, while the pool it waits on is the real constraint.
Trade-offs
Attacking per-request overhead works and it is architectural: batching changes an API contract, removing a middleware layer removes a cross-cutting guarantee with it, dropping telemetry trades future debugging for present latency. These are not local refactors and need a sponsor.
The cheap alternative is honest and unpopular: accept the per-request cost and add instances. If the profile is genuinely flat and CPU accounts for the latency, the code may already be efficient, and this is a capacity decision priced against the engineering cost of removing a layer.
When not to use it
Do not reach for this interpretation before checking the measurement. A 30-second profile of a low-traffic service is flat for lack of samples, not because cost is distributed; an hour at 99 Hz is hundreds of thousands of samples and can be trusted.
And once a profile does show a 40% frame, stop reasoning about distribution and optimise it: this is a diagnosis for one shape of evidence, not an argument against optimisation.
Interview question
Q: "Your service misses a 300 ms p99. The CPU profile has no frame above 4%. Someone proposes optimising the top frame and someone else proposes raising the sampling rate. What do you do?"
What a strong answer covers: computing CPU milliseconds per request and comparing it with wall-clock latency before anything else; naming the off-CPU candidates and the tools that find them; Amdahl's ceiling on the 4% frame as a quantified rejection of the first proposal; why more samples cannot create concentration that is absent; and, if on-CPU and flat, proposing a reduction in units of work with its contract costs stated.
Quick check
Quiz: 12 ms of CPU per request and a 300 ms p99. How much of the latency can a CPU profile explain? About 4% — the rest is off CPU, so the investigation is locks, IO, pools or GC.
Flashcard: What does a flame graph with no frame above 4% tell you to do? — Stop hunting for a hot function: cut the number of requests and the per-request path, or measure off-CPU time, after checking that symbolisation is not the reason it looks flat.