concept

Metric Staleness

also called Staleness Marker, Stale Series

The rules deciding how long a time series keeps answering queries after it stops being written - which determines whether a dead process's last value is served as current, and whether an alert on a vanished series fires or goes quiet.

metricsprometheusalertinglookbackno-data

A pod is deleted at 12:00. Its queue-depth gauge last reported 8400, above the alerting threshold. At 12:01 an instant query still returns 8400. At 12:06 the series is gone and the rule's result is empty, which most engines render as healthy. One alert fired about a process that no longer existed, then stopped firing without anything being fixed.

Metric staleness is the set of rules behind that sequence: how far back a query reaches for the newest sample, and what happens when a writer stops. Get it wrong and you have two symmetrical failures — stale values served as current, and disappearance served as health.

Why it matters

"No data" is the most common silent failure in alerting. A renamed metric, a relabelled job, a scaled-to-zero deployment and a crashed exporter all produce the same query result as a healthy system: nothing. The alert that protected a service has been deleted by a refactor, with no error to say so.

The mirror failure costs in the other direction: a capacity decision taken from a gauge whose writer died five minutes ago is a decision about a system that no longer exists, and nothing on the graph marks the line as a held value.

Implementation patterns

  • Explicit staleness markers on the pull path. Prometheus writes a special NaN marker when a target disappears or a scrape stops exposing a series, so instant selectors stop returning it at the marker rather than at the end of the lookback window.
  • A bounded lookback for everything else. --query.lookback-delta, five minutes by default, is how far an instant query reaches; push paths carry no markers, so the lookback alone ends the series.
  • Absence as a first-class alert. absent_over_time(series[10m]) on the inputs of every critical rule, so disappearance fires instead of passing.
  • Freshness as data. Emit a last_updated_seconds gauge next to slow or pushed metrics so a panel can show the age of the number.

Industry example

The behaviour is specified rather than emergent: Prometheus 2.0's staleness handling was presented at PromCon in 2017 and its rules are documented. The asymmetry it creates is what teams meet in production. Scraped series end cleanly; series arriving by remote write, a push gateway or an OTLP push path get no markers, so their last value is served for the full lookback window. Grafana's Tempo project carries an open request to emit staleness markers over remote write, which says the gap is structural rather than a local misconfiguration. An estate mixing scraped and pushed metrics has two staleness behaviours on one dashboard, and the panels do not say which is which.

Failure scenarios

  • The refactor that deletes an alert. A metric is renamed, the rule matches nothing, and it sits permanently non-firing with no error.
  • Scale-to-zero read as an outage. A deployment scaled to zero overnight makes its series absent, and an absence alert written without that in mind pages every night.

Trade-offs

Shortening the lookback delta makes staleness honest, because a value stops being served soon after its writer stops, and it makes every metric written less often than the window vanish between samples. Lengthening it keeps sparse series continuous and lets dead processes' values persist, the more dangerous error because it is invisible.

The bill is paid by alerting, not by dashboards. A human reading a graph can be told the line is old; a threshold comparison cannot. That asymmetry is the argument for spending effort on absence guards and freshness gauges rather than on tuning the window.

When not to use it

Do not put absence alerts on every series. On a fleet with high churn — short-lived pods, per-build labels, autoscaled workers — series appear and disappear by design, so absence alerts there are pure noise. Guard the series an alert depends on, not the series that exist. And when a metric is written every few seconds by a long-lived scraped target, default staleness handling is already correct.

Interview question

Q: An alert on queue_depth > 1000 has not fired in three months and the queue has definitely been deep. Walk me through everything that could be true, and what you would add so this class of problem announces itself.

What a strong answer covers: that an empty result cannot cross a threshold, so a renamed metric, a changed label or a dropped scrape job all produce silence identical to health; that pushed series get no marker and so behave differently from scraped ones; the guards — absent_over_time on the rule's inputs, rule unit tests against recorded data, a freshness gauge for pushed metrics, and a dead man's switch proving the delivery path; and the judgement to guard the few series alerts depend on rather than a churning fleet.

Quick check

Quiz: Why does deleting a pod end its scraped series immediately but leave its pushed series answering queries for five minutes? The scrape path writes an explicit NaN staleness marker when a series stops being exposed; push and remote-write paths carry no marker, so only the query engine's lookback delta ends the series.

Flashcard: An alert rule's query returns no series. Does it fire? — No, and that is the failure: no samples means no comparison, which renders as healthy. Guard every critical rule with an absence check.