concept

Push Token Decay

also called Stale Registration Token, Push Target Rot

The steady accumulation of push targets that no longer reach a device - which inflates fan-out cost and makes delivery metrics describe a list rather than an audience.

pushfcmapnsfan-outmetrics

A product manager reports 2.1 million devices reachable by push. The installed base is 1.3 million. Nobody lied: the fan-out list has been accumulating registration tokens for three years, and a token whose device is long gone does not announce itself. The gap between the two numbers is the error bar on every delivery statistic the team has.

Decay has two sources. Tokens become permanently invalid when an app is uninstalled, data is cleared, or the instance is restored to another device. And tokens stay syntactically valid while the device has simply stopped existing, because nothing in the push system is obliged to tell you about silence.

Why it matters

A rotten list corrupts the only metric most teams have. "Delivered" measured against a list that is 30 to 40% dead is not a measurement of anything, and it is exactly the number quoted in an incident review when users report stale data. The fan-out cost is also real: duplicated sends, wasted queue capacity and per-message charges on third-party delivery platforms.

The subtler cost is correctness. If push is the trigger for a background sync and the token is dead, the device never syncs on time and the server has no indication, so the failure is invisible on both sides.

Implementation patterns

  • Delete on permanent errors. FCM's HTTP v1 API returns 404 UNREGISTERED for a token that will never be valid again; delete the row rather than retrying. INVALID_ARGUMENT (400) signals a malformed token, which is a bug in your own storage.
  • Re-register on every app launch and store a last-seen timestamp next to each token, because an absent device produces no error and only a timestamp reveals it.
  • Expire by age. Firebase treats a registration whose app instance has not connected for a month as stale, and since the change announced in April 2024 and enforced from 15 May 2024, Android tokens are expired after 270 days of inactivity. Pick a threshold in that range and apply it.
  • One row per installation, not per user. A user with three devices has three tokens, and a token that moves to a new installation must replace the old row rather than add one.
  • Measure outcomes, not sends. The share of devices whose last successful sync is under an hour is a real number; delivery receipts are not.
  • Keep a scheduled fetch as the floor. Push is best-effort by design, so correctness must not depend on it.

Industry example

The two platform services document the behaviour precisely. FCM's registration-management guidance tells senders to delete tokens on a 404 and to refresh stored tokens periodically, and treats a month of inactivity as staleness. Apple's APNs answers a send with HTTP status 410 and the reason Unregistered when a token is no longer active for the topic, with a timestamp for when it was last confirmed invalid, and instructs senders to stop until a token with a later timestamp is registered. Both make the same assumption: the sender owns the list and is expected to prune it. No platform prunes it for you, and both decline to distinguish "device off" from "device gone" in the send response, which is why age-based expiry is not optional.

Failure scenarios

  • A fan-out list that grows monotonically, so cost scales with history instead of with the audience.
  • Delivery dashboards quoted as evidence during an incident, hiding a cohort that has not synced in days.
  • A token reassigned to a different installation after a device restore, sending one user's notifications to another, which is a privacy incident and not merely a bug.
  • Retrying 404s forever in a generic retry wrapper, turning dead rows into permanent background load.
  • Per-user storage of a single token, so a user with two phones only ever receives on whichever registered last.

Trade-offs

Pruning aggressively risks removing a token for a device that is merely off for a long holiday, and that device will re-register on next launch at no cost. Pruning conservatively keeps the cost and keeps the metric wrong. The asymmetry favours pruning: the recovery path is automatic, and the cost of a wrong metric is paid during incidents.

The other trade-off is where staleness is tracked. Last-seen timestamps mean every launch writes to the token store, which at tens of millions of installations is a non-trivial write rate, usually handled by batching and by accepting approximate timestamps.

When not to use it

A small application sending occasional broadcast notifications to a few thousand devices does not need an expiry regime; deleting on permanent errors is enough. The regime starts earning its complexity when push is a functional dependency — a sync trigger, a dispatch channel, a two-factor prompt — or when the fan-out list is large enough for its cost to be visible.

Interview question

Q: Your push provider reports 99.3% delivered and support says a few per cent of couriers see hours-old job lists. Where do you look, and what would you change about what you measure?

What a strong answer covers: that "delivered" means accepted by the platform for a token, not shown on a device; that the list itself may be substantially dead; the pruning rules on permanent errors and by age; and the replacement metric, which is the share of devices whose last successful sync is recent, with a server-side scheduled fetch as the floor beneath push.

Quick check

Quiz: Which push response must result in deleting the stored token? Answer: a 404 UNREGISTERED from FCM; that token will never be valid again, so retrying is pure waste.

Flashcard: Why does a push list rot without producing any errors? — A device that has simply stopped existing leaves a syntactically valid token, so only a last-seen timestamp plus age-based expiry removes it.