Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
3 to work through
-
intermediate Multiple choice
When does APM tell you something that distributed tracing and metrics cannot?
2 min answer -
advanced
In a multi-tenant platform, aggregate application performance metrics look healthy while specific tenants experience severe slowness. What must performance monitoring do differently?
2 min answer -
advanced
Review this instrumentation. A checkout service emits about 400 spans per request, every span carries roughly 30 attributes including the full request body, tracing runs at 100% with no sampling, and the team reports that they cannot find anything in the traces. What would you remove, what would you change, and what would you keep even though it looks excessive?
3 min answer
3 terms in this topic
APM Transaction Tracing
Instrumentation that attributes application latency and errors to specific code paths, database queries and external calls, usually with automatic tracing.
toolApplication Performance Monitoring
Instrumentation that attributes latency and errors to code paths and dependencies inside a service — the layer between metrics and a profiler.
practiceOutlier Alerting
Alerting on the worst-affected segment rather than on an aggregate, because in skewed populations the aggregate is dominated by the segment that is fine.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.