Dead Letter Queue
also called DLQ, Poison Message Queue
A separate destination for messages that cannot be processed successfully, so one bad message does not halt the stream behind it.
Without one, a single unprocessable message — malformed payload, a schema the consumer does not understand, a reference to a record that no longer exists — blocks its partition permanently. The consumer retries, fails, retries, and everything behind it waits. This is the classic poison message outage, and it takes down a pipeline for hours while someone works out which offset to skip.
The pattern is to move the message aside after a bounded number of attempts, along with the metadata needed to understand it later: the original topic and offset, the failure reason, the stack trace, the consumer version and the timestamp. Without that context the queue is a directory of unexplainable payloads.
The part that fails in practice is everything after the message is moved. A dead letter queue with no alert, no owner and no replay tooling is a data loss mechanism with good branding — messages accumulate silently, and the first anyone knows is a reconciliation gap months later.
The operational requirements are therefore explicit: alert on any arrival, since a healthy stream should produce none; a supported replay path once the defect is fixed; and a retention policy that does not quietly delete unresolved records. Distinguish transient failures, which belong in retry with backoff, from permanent ones, which belong here immediately.