Symptom-Based Alerting
also called Alert on Effects, User-Visible Alerting
Paging on what the user experiences and investigating causes with dashboards - because cause-based alerts fire in clusters during a single incident and bury the signal they exist to provide.
An alert on a cause — high CPU, a deep queue, a saturated pool, a dropped cache hit rate — fires whenever that condition occurs, whether or not anything is wrong, and fires alongside a dozen others during a single incident. One event produces forty pages, and tuning thresholds does not change that because the problem is structural.
Symptom-based alerting pages on the effect: the success rate of a user journey, its latency, its volume against baseline. One incident produces one page.
Why it matters
Alert volume is the primary determinant of whether alerts are read. A rotation receiving hundreds of pages will miss the one that mattered, and no amount of individual alert quality compensates for the aggregate.
It also produces better diagnosis, because the responder starts from "what is the user experiencing" rather than from whichever cause happened to fire first — which is frequently a symptom of the real cause rather than the cause itself.
Implementation patterns
- The test for every alert: if this fires at 3am, is there something a human must do right now? If not, it is a dashboard entry or a ticket. Applying this honestly typically removes most of the volume.
- Keep causes on dashboards, arranged so the responder opens one page after being paged and sees every contributing signal.
- Delete alerts on individual instances in fleets where instance loss is handled automatically.
- Delete alerts on utilisation rather than on the effect of utilisation. High CPU is only a problem if something is degraded.
- Deduplicate across layers — the same failure detected by infrastructure, application and synthetic monitoring produces three pages for one event.
- Measure the percentage of pages that resulted in an action, and review monthly. Alerts that have never produced an action are the deletion list.
- Measure pages per shift against a number the team has agreed is sustainable, since a target nobody set will not be met.
Industry example
Marketplace and mobility platforms such as Ola and Swiggy have deep dependency graphs where a single upstream degradation produces alerts at every layer. The improvement that changes on-call sustainability is not better thresholds but collapsing to a small number of journey-level symptom alerts — trip assignment success, order placement success, payment completion — with everything else demoted to dashboards.
Failure scenarios
- Cause-based alerting, producing clusters per incident.
- Alerts retained because deleting them feels risky, which is why most survive.
- Symptom alerts with no dashboard behind them, leaving the responder paged with nowhere to look.
- Symptom thresholds set against a fixed number rather than a seasonal baseline, so they either never fire or fire constantly.
- No measurement of actionability, so the deletion decision has no evidence and becomes a debate.
Trade-offs
Symptom alerting detects later than cause alerting in some cases: a queue growing is visible before the user experiences a delay. That earlier warning has genuine value for slow-building problems.
The resolution is not to reinstate cause pages but to use leading indicators as tickets or low-severity notifications rather than pages, and to reserve the pager for user-visible effect. The earlier signal is still available to anyone looking; it simply does not wake anyone.
Interview question
"You have 200 alerts and want to get to 20. Describe your method for deciding which survive, and tell me what you would do about the one everyone insists must stay even though it has never been acted on."