Differential Profiling
also called Profile Diffing, Comparative Profiling
Comparing two normalised profiles of one workload to attribute a performance change to a specific code path - the step that turns "release 2.4 is slower" into a named function, and the normalisation errors that make the comparison lie.
Release 2.4 uses 18% more CPU than 2.3 at the same request rate, and both profiles are available. Reading them side by side almost never works: the eye cannot spot a three-percentage-point shift in a tower of frames, and the two captures differ in a dozen irrelevant ways.
Differential profiling makes the comparison arithmetic instead of visual: normalise both profiles to a common denominator, subtract the stack trees, and read the residual. Rendered as a differential flame graph, growth and shrinkage are coloured and unchanged frames drop out of attention, attributing the change to a code path rather than to a service.
Why it matters
Most regressions are found by a metric and diagnosed by guesswork: a CPU or latency graph localises a change to a service and a time, and says nothing about which code moved. Bisecting by reverting commits is the usual fallback, and it costs a deploy per hypothesis on a system that may not reproduce the regression outside production.
Done carelessly the comparison produces confident wrong answers: each normalisation mistake generates a plausible culprit, and a team acting on one spends a sprint optimising a function that was never the problem.
Implementation patterns
- Normalise before subtracting. Divide sample counts by requests served, because raw counts scale with duration and sample rate.
- Capture under matched load. A canary beside the stable version behind one load balancer is cleanest, since both see the same request mix.
- Keep symbol files for every build, which makes symbol retention a release-process requirement rather than a tooling preference.
- Diff the right resource. If wall-clock time dominates CPU time, diff off-CPU profiles instead, where lock contention and blocking calls live.
Industry example
Discord's 2020 post on rewriting its Read States service from Go to Rust is a before-and-after comparison of one workload, and its reasoning is what a differential view formalises. The service showed latency spikes on a two-minute cadence matching Go's forced garbage-collection interval, in a service that allocated very little, pointing at the collector walking a large live heap rather than at garbage production or request volume. The decisive evidence was that the period matched a runtime interval and was independent of load — a comparison across time inside one version, before any comparison across versions. The Rust implementation did not show the spikes.
Two limits matter for anyone reaching for that story. A CPU profile averaged over a minute would likely have missed it, because collector work is a thin slice of total samples. And a language rewrite is the most expensive response available to a tail-latency graph, rarely defensible outside a service that walks a large mostly-static cache.
Failure scenarios
- Unmatched traffic mix. The newer profile was captured during a nightly batch window, so the diff shows the batch and the team optimises a report generator.
- Inlining differences. A compiler inlines a function in one build and not the other, so a frame vanishes with no behaviour change and the diff blames its caller.
- Missing frame pointers. Stacks truncate and work lands on the wrong ancestor; distribution-built binaries commonly omit them.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Differential profiling | Attribution to a code path; no hypothesis needed | Two comparable captures; symbol retention; normalisation care |
| Bisect by revert | No profiling infrastructure | A deploy per hypothesis; days elapsed; needs a reproduction |
The decision rule is the size of the change relative to the baseline. A large absolute share of CPU shows up in a single profile; a few percentage points needs the diff — and across a large fleet, a few percentage points is where the money is.
When not to use it
If the regression is 18% of CPU on a service that is 2% of the fleet, note it and spend the effort on something in the top ten: the diff is cheap and acting on it is not. If the versions cannot run under the same load, a staged canary comes first. And when the symptom is latency at low CPU utilisation, compare off-CPU time or span durations instead, because a CPU diff there is an honest answer to the wrong question.
Interview question
Q: A canary at 5% traffic uses 18% more CPU per request than stable, and you have continuous profiling on both. Localise it, and give me three ways the comparison could be wrong.
What a strong answer covers: normalising per request rather than per capture; using the canary because both versions see the same traffic mix; subtracting stack trees and reading the residual rather than eyeballing two flame graphs; the failure list — mismatched load windows, unnormalised sample rates, inlining differences that rename frames, and missing frame pointers truncating stacks; checking CPU against wall-clock time first; and the judgement that an 18% regression on a small service may not be worth acting on.
Quick check
Quiz: You diff a 60-second profile of A against a 300-second profile of B and everything in B looks larger. What went wrong? Counts were not normalised; divide by requests served, because raw counts scale with capture duration and sample rate.
Flashcard: Two profiles of one service, one per release. What comes before subtracting them? — Normalise both, ideally to samples per request, and confirm both captures saw the same traffic mix.