metric

Dead-Letter Budget

also called Diversion Budget, Poison Rate Budget

A declared ceiling on the share and the age of records a pipeline may divert before it counts as failing rather than coping - with an owner and a defined action attached to each breach.

dead letterslosdata qualityerror budgetownership

Forty thousand records sit in a dead-letter topic, the oldest from seven months ago, and nobody knew. Every individual decision that produced that state was defensible: divert the record, continue processing, alert if the topic is non-empty, silence the alert because it is always non-empty.

A dead-letter budget replaces "is the topic empty" with two numbers that can actually be operated: the maximum fraction of records that may be diverted in a window, and the maximum age of the oldest undiverted-and-unresolved record. Both have an owner, and breaching either one triggers a named action.

Why it matters

A dead-letter queue converts an availability problem into a data-quality problem. That is a good trade, and it is only a trade if someone is accounting for the data quality side. Without a budget, diversion is unbounded: a schema change that breaks 4% of records looks identical in every dashboard to a bad record once a week, because both produce a non-empty topic and neither produces an error.

The budget also makes a governance question answerable. When a downstream report is questioned, "the pipeline was inside its 0.01% diversion budget and the oldest unresolved record was under 24 hours" is evidence. "The dead-letter topic had some records in it" is not.

Implementation patterns

  • Express the ceiling as a rate, not a count. 40 thousand records is meaningless without the denominator; 0.01% of a 200-million-record day is 20 thousand.
  • Publish age of oldest unresolved record alongside the rate. Rate catches the schema change; age catches the slow leak that never trips a rate threshold.
  • Attach an action per breach. Rate breach: page the producing team, because a rate breach is a systematic fault. Age breach: an item in the owning team's queue with a date, because an age breach is neglect rather than an incident.
  • Make the records judgeable. Every diverted record carries its schema id and the code version that rejected it — Kafka Connect's context headers give topic, partition, offset, stage and exception when errors.deadletterqueue.context.headers.enable is on. Without the version, nobody can tell whether a record is still invalid.
  • Set the budget to zero where records are order-dependent. A ledger stream's correct budget is no diversions at all; halt instead.

Industry example

Kafka Connect ships the honest default, and it reads differently once a pipeline has been in production for a year: errors.tolerance is none, so a connector fails on the first bad record and the dead-letter topic is empty unless errors.tolerance is set to all and a topic name configured. Teams flip that pair to stop pages, and in doing so change the failure mode from loud to silent without changing any accounting. The setting is one line; the budget is the missing half of the change.

Failure scenarios

  • The graveyard: months of records nobody can interpret, because the code that rejected them has changed a dozen times.
  • The silent schema break: 4% of records diverted for a week, and every downstream aggregate quietly 4% low.
  • Out-of-order replay: a bulk replay of diverted records applied after later records, which corrupts order-dependent state while looking like remediation.
  • Ownership gap: the platform team is paged for records whose meaning only the producing team understands, so the page is acknowledged and nothing happens.

Trade-offs

A budget costs accounting work: a denominator, a dashboard, an owner per topic, and a review cadence. It also forces an uncomfortable conversation, because agreeing a non-zero budget means stating out loud that some records will be lost. The alternative is not zero loss — it is the same loss, unmeasured.

When not to use it

Do not put a budget on a stream where diverting even one record is wrong. Order-dependent folds (balances, inventory, state machines) should halt, and dressing that up as a budget of zero adds ceremony to a decision that is really about partitioning the blast radius so only the affected key stops. Equally, for a pipeline with two consumers and a thousand records a day, a weekly look at the topic is proportionate and a formal budget is theatre.

Interview question

Q: You inherit a pipeline whose dead-letter topic holds 40 thousand records over seven months. What do you do in the first week, and what do you put in place so it cannot recur?

What a strong answer covers: classify before replaying — which records are still invalid, which were rejected by a bug since fixed, which are duplicates of data that arrived another way · decide per class whether replay is safe given ordering · then the standing controls: a rate budget with a denominator, an age-of-oldest metric, a named owner, the schema id and code version on every record, and an explicit statement of which streams may not divert at all.

Quick check

Quiz: Why is "alert when the dead-letter topic is non-empty" a weak control? It has no denominator and no age, so a systematic 4% diversion and one bad record a week produce the same signal, and the alert is silenced within a fortnight.

Flashcard: What two numbers make up a dead-letter budget? — The maximum share of records that may be diverted in a window and the maximum age of the oldest unresolved record, each with an owner and a defined action.