Profiling
Continuous CPU and memory attribution in production.
4 to work through
-
advanced
A data platform has good tracing and still cannot explain why a specific job is slow. What does tracing not tell you, and what does?
2 min answer -
advanced Multiple choice
A developer-tools company needs to find a performance regression that only appears under real production workloads. What are the options for profiling in production, and what are their costs?
2 min answer -
advanced
A service is slow and CPU utilisation is 4%. What do you profile and what do you expect to find?
2 min answer -
advanced
Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.
3 min answer
3 terms in this topic
Continuous Profiling
Sampling CPU, memory and lock profiles from production continuously at low overhead, so resource usage can be attributed to specific code paths.
practiceDifferential Profiling
Comparing two normalised profiles of one workload to attribute a performance change to a specific code path - the step that turns "release 2.4 is slo…
practiceProfiling
Attributing resource consumption to specific code paths — the tool for "why is this slow" once you know where.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.