Alert Fatigue
How noise makes the real page invisible, and the structural fix.
6 to work through
-
intermediate
A team has 340 alerts. On-call receives roughly 40 pages a week and most are ignored. What would you change?
2 min answer -
intermediate Multiple choice
A team receives hundreds of alerts a day and misses a genuine outage. What is the fastest structural fix, and what should replace the current approach?
2 min answer -
intermediate
An on-call engineer receives so many alerts that they routinely acknowledge without reading. What has failed, and what is the risk beyond tiredness?
2 min answer -
intermediate
An on-call team receives 60 alerts per shift and has stopped reading most of them. What is the structural fix, and what should the alerting rules actually be based on?
3 min answer -
intermediate
Half your pages result in no action being taken. How do you fix that without reducing coverage?
2 min answer -
advanced
Your team receives 200 pages a week and the on-call rotation has lost two engineers in six months. Fix it.
2 min answer
3 terms in this topic
Alert Actionability
The proportion of alerts that result in a human taking action, used as the primary quality measure of an alerting system.
conceptAlert Fatigue in Practice
The state in which alerts are ignored because most of them do not matter — a reliability failure caused by monitoring rather than prevented by it.
conceptSymptom-Based Alerting
Paging on what the user experiences and investigating causes with dashboards - because cause-based alerts fire in clusters during a single incident a…
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.