beginner 2 min answer

Your team's on-call phone has not been paged in nine days and every dashboard is green. Why is that not yet evidence that anything is healthy, and what single mechanism turns silence into a signal?

alertingdead-mans-switchwatchdogheartbeatself-monitoring
Show the full answer Hide the answer

The mechanism

An alerting pipeline has four links: collection, rule evaluation, routing, and notification delivery. A break in any one of them produces exactly the observation you are looking at — no page. The failure mode of an alerting system is indistinguishable from its success, so silence carries no information at all until you deliberately make it carry some.

The concrete ways the chain breaks are mundane. A scrape target's labels change and the rule's query matches nothing. Someone leaves a silence in place after an incident. A routing tree sends one team's alerts to a receiver whose credential expired. The evaluation process runs out of memory and restarts in a loop. None of these emits an error to anyone who would notice.

What a dead man's switch asserts

One rule with a condition that is always true — conventionally named Watchdog in the Prometheus operator bundles used in production almost everywhere self-hosted — routed continuously to a heartbeat service outside the monitored system. That external service pages when the pings stop.

The assertion is narrow and worth stating precisely: an alert generated inside the platform reached a system outside it within the last few minutes. It says nothing about whether your rules are correct. It says everything about whether a correct rule could have reached you.

How to size it

Send the heartbeat every minute and set the external timeout to at least three intervals, commonly 5 minutes, because a single dropped ping should not page anyone. The notification for the heartbeat must travel through a provider and a channel the monitored platform does not depend on. If the heartbeat's own page would be delivered by the system it is watching, you have rebuilt the coupling and bought nothing.

What it does not cover, and what does

  • A rule whose query returns empty because a metric was renamed: guard each critical alert with absent_over_time() on the series it depends on, so absence fires instead of passing.
  • Alerts that exist but are wrong: unit-test rules against recorded time series in CI.
  • A broken route for one team: fire a real test alert through every route on a schedule and require someone to acknowledge it.
  • Silences that outlive their incident: alert on any silence older than 24 hours.

When not to bother

If alerting is delivered by a managed provider that pages you on its own delivery failure through an independent path, and you have tested that claim rather than read it in a datasheet, the switch adds a dependency for little gain. On any self-hosted stack the cost is one rule and one free external check, which makes it the cheapest reliability control available and leaves no honest argument for skipping it.