advanced 3 min answer

Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.

discordprofilinggarbage-collectiontail-latencyplatform-cost
Show the full answer Hide the answer

What the interviewer is testing

Whether you can separate three diagnostics that people blur together, and whether you can price a platform decision rather than approve a tool. Tracing answers where the time went between services. A CPU profile answers which code is hot. Runtime telemetry answers why the process was not running at all. The Discord case is the third question, and a CPU profile alone would probably have missed it: garbage-collection work averaged over a minute of samples is a thin sliver, while the user-visible damage is concentrated in a few hundred milliseconds.

The clarifying questions that change the answer

Is the spike periodic or load-correlated? Periodic and independent of traffic points at the runtime or at a scheduled job, not at request handling. Does it appear in the service's own latency or only at the caller? Only at the caller means queueing, scheduling or the network. Is CPU time elevated during the spike, or is the process simply not executing? Those are opposite investigations. Discord's evidence was that the period matched Go's minimum forced collection interval, in a service whose steady-state allocation was low — which pointed at the collector's scan of a large live heap rather than at garbage production.

What I would fund

A sampling CPU profiler on by default across the fleet at 19 to 99 Hz per core, plus per-process runtime pause metrics — collection count, total pause time, heap live bytes — which cost almost nothing and are the signal that actually diagnoses this class. An eBPF-based agent at that rate typically costs on the order of 1% of a core and a couple of hundred MB of memory per node; across 800 nodes that is under 1% of fleet compute and roughly 160 GB of agent memory. The storage side is the part people mis-plan: profile cost scales with distinct stacks multiplied by label sets multiplied by retention, not with request volume, so a fleet-wide rollout with a pod label is the same mistake as a pod label on a metric.

What I would refuse

Heap profiling at every allocation, which changes the program's performance enough to invalidate what it measures. Per-request profiling, which is tracing wearing a profiler's clothes. Symbolication of stripped binaries at query time on a hot path. And a fleet-wide rollout before one team has used it to close one real regression, because an always-on tool nobody reads is a recurring bill with no payer.

Common weak answers

"Turn on the profiler and look at the flame graph" ignores that a flame graph aggregates samples and dissolves a periodic 200 ms event into background noise; you need the profile sliced to the spike window, or an off-CPU view, before it says anything. "Add more replicas" treats a per-process pause as a capacity problem: every replica pauses on its own schedule, so the fleet-wide probability that some replica is paused goes up, not down.

When always-on profiling is the wrong spend

If nobody on the team has read a profile in the last quarter, buy the runtime pause metrics and the ability to capture a profile on demand, and stop there. And the senior layer on Discord specifically: a language rewrite is the last move, not the first. They had a service whose work was walking a large mostly-static cache, which is close to the worst case for a tracing collector. For most services the answer is reducing live heap, tuning collector parameters, or moving the cache out of the managed heap. A four-engineer team responding to a tail-latency graph with a rewrite is choosing the most expensive available option.