concept

Metric Temporality

also called Delta Temporality, Cumulative Temporality

Whether a reported metric point carries a running total since a start time or only the change during one interval - the choice that decides whether process restarts or duplicate deliveries are what corrupts your rate graphs.

opentelemetrymetricscountersexactly-oncecollector

A counter reports 4000. A minute later it reports 4900. Whether 900 events happened or 4900 did depends on a property of the stream that no dashboard displays. Cumulative temporality means each point is a running total since an explicit start time; delta temporality means each point is the count for that interval alone. Both are in wide production use.

Cumulative pushes the hard problem onto the reader, which must treat a decrease as a process restart rather than negative work and know the new start time. Delta pushes it onto the transport: with no running total to re-synchronise against, a duplicated batch over-counts and a dropped batch under-counts permanently.

Why it matters

A temporality mismatch produces graphs that are wrong rather than empty, which is the expensive failure. A missing series gets noticed; a rate understated by 20% because reset detection never ran gets used to size a fleet. OpenTelemetry's data model carries temporality on the wire because it cannot be inferred from the values.

It also decides where state lives: a delta pipeline serving a pull-based backend needs something to accumulate, and if that something is the collector, the collector is now stateful and its restart is a data event.

Implementation patterns

  • Set it once, at the SDK, to match the backend. Pull-based systems and remote write want cumulative; several vendor ingest paths want delta.
  • Always emit start_time_unix_nano with cumulative points, because a backend guessing from a decrease misses a reset that happens to land on a higher value.
  • Convert in one place. The collector's cumulativetodelta processor handles one direction; the reverse needs per-series accumulation with the memory and restart consequences that implies.
  • Never mix temporalities under one metric name. The aggregate is nonsense and still draws a plausible line.

Industry example

OpenTelemetry's Prometheus and OpenMetrics compatibility specification sets out the mapping in both directions, including that exponential histograms convert one-to-one onto Prometheus native histograms, stable since Prometheus 3.8. Vendor preferences diverge for a structural reason: delta lets ingestion be a pure append with no per-series read, the same reasoning behind Datadog's Husky design in 2022 where writers never read what they write into. The consequence for a team standardising in 2026 is that this decision is made once and then constrains which backends can be added later without a conversion layer.

Failure scenarios

  • Reset blindness. A pod restarts, a cumulative counter drops from 1.2 million to 300, and a backend without reset handling draws a negative spike or clamps to zero, losing that traffic.
  • Delta double-count on retry. The exporter times out after the backend committed, retries, and the interval is applied twice. The graph is 2x high for a minute and nothing errors.
  • Silent under-count. A dropped delta batch is gone; no later point carries the missing work, so a daily total is quietly low forever.
  • Stateful collector restart. A collector accumulating delta into cumulative restarts and emits a reset to every downstream series at once.

Trade-offs

Choose Gains Pays
Cumulative Restart-tolerant; a late or duplicate point is harmless because it is a total Reader must detect resets; needs start timestamps; useless for very short-lived series
Delta No reset problem; cheap append-only ingestion; suits elastic workloads Needs exactly-once delivery to be exact; a dropped interval is unrecoverable

The decision rule follows process lifetime. Long-lived services behind a pull-based backend belong on cumulative; workloads where a series may exist for 40 seconds belong on delta, because a cumulative counter that starts and dies inside one scrape interval is never observed at all.

When not to use it

Temporality applies to sums, counters and histograms, not to gauges, which carry a value at an instant; forcing a gauge through a delta conversion produces a plausible derivative of a level and means nothing. If the estate is all long-lived services behind one pull-based backend, this is a default to leave alone. The analysis earns its time when there are several backends, elastic workloads or a migration in progress.

Interview question

Q: You are standardising 60 services on OpenTelemetry metrics. One team runs 30-second batch jobs in a serverless runtime and the platform backend is pull-based. What temporality do you set, for whom, and what does the mismatch cost you?

What a strong answer covers: cumulative for the long-lived services because the pull-based backend expects it and reset handling is solved there; delta for the 30-second jobs because a cumulative counter whose process dies inside a scrape interval is never collected; a collector accumulating those deltas into cumulative, with the admission that this makes the collector stateful and its restart a data event; and the rule that one metric name never carries two temporalities.

Quick check

Quiz: A cumulative counter's value drops after a deploy. What must the backend do? Read the decrease as a process restart rather than negative work, using the point's start timestamp to establish the new origin; without it, a reset landing on a higher value is missed.

Flashcard: Delta or cumulative for a 30-second serverless job? — Delta. A cumulative counter that starts and dies inside one scrape interval is never collected, so the work is invisible; the price is that a dropped delta batch is lost permanently rather than recoverable from a running total.