Compensation Deadline
also called Undo Deadline, Compensation Escalation Window
A bounded time within which a saga's undo actions must succeed before the saga is declared unrecoverable and escalated to a person - turning a silent inconsistency into a queue somebody owns.
A booking saga captures a card payment, reserves a room, then fails at the loyalty step. The orchestrator starts compensating: refund the payment. The payment gateway returns 503 for the next 40 minutes, so the refund retries with backoff. It is still retrying three days later, because the card scheme's refund window for that transaction closed at 48 hours and every attempt now returns a permanent error that the retry policy treats as transient.
Nothing is broken, in the sense the dashboards use. The saga is "in progress". The money is with the merchant, the room is free, and the customer has been charged for a booking that does not exist. The design assumed compensation eventually succeeds, and there is no state for "it will not".
A compensation deadline is the missing state. Each compensating action gets a wall-clock bound, derived from the business constraint it depends on, after which the saga stops retrying and becomes a work item with an owner, an amount at risk and a customer attached.
Why it matters
A saga buys atomicity by substituting compensation for rollback, and the 1987 Sagas paper that introduced the idea assumed compensating transactions always complete. In production they do not: refund windows close, inventory is resold, the confirmation email has already been read, a downstream partner retires the endpoint that performed the undo.
The arithmetic makes the exposure concrete. A platform running 100,000 sagas a day where 0.4% need compensation and 2% of those cannot complete produces about 8 permanently stuck sagas a day. With no terminal state they accumulate invisibly, and the business learns the number from a month-end reconciliation rather than from a queue. Unbounded retry is how a consistency problem becomes a finance problem.
Implementation patterns
- Derive each deadline from the external constraint, and set it shorter. If the refund window is 48 hours, the compensation deadline is 24, leaving time for a human attempt.
- A terminal
compensation_failedstate, distinct from "retrying", with the failing step, the last error and the elapsed time recorded. - Classify errors before retrying. A permanent rejection must not be retried on the transient policy; the bug above is a retry loop treating a closed window as a 503.
- Escalate to a durable queue, not an alert. The item carries saga id, step, amount at risk and the customer, so it can be worked and prioritised by value rather than by arrival order.
- Idempotency keys on every compensating action, so a human pressing retry cannot double-refund.
- An amount-at-risk gauge as the service level indicator. The count of stuck sagas understates a single large one.
Industry example
The pattern is visible in any marketplace that sells third-party inventory in production: the forward steps are owned by the platform, and at least one compensation depends on a partner's willingness to accept it. Platforms that run this well expose the stuck-saga queue to operations as a first-class work type, staffed like any other exception queue, with a target age rather than a target of zero. The ones that do not run it well discover the backlog when a regulator or an auditor asks what happens to a failed booking.
Failure scenarios
- Infinite retry of a permanent error, so the saga never reaches a human.
- Compensation that succeeds partially — the refund is issued, the loyalty points are not reversed — with no record of which half completed.
- An escalation path that is a page, which gets acknowledged and forgotten, instead of a queue with an age metric.
- A deadline longer than the external window, which guarantees the human attempt also fails.
- Compensating a step that never ran, because the saga log recorded the outcome rather than the intent.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Deadline plus escalation | Bounded inconsistency; a number you can report | An exception queue to staff, and ops tooling to work it |
| Unbounded retry | No operational surface | Unbounded, invisible inconsistency |
The real cost is organisational: a compensation deadline creates work that lands on people, which is why teams avoid defining one. The work exists either way; the deadline decides whether it is visible.
When not to use it
When every compensating action is internal and cannot fail permanently — reversing a row in your own database, releasing a reservation your own service holds. There, retry with backoff genuinely does converge, and a deadline adds a state nobody will ever use. It flips as soon as one compensation leaves your trust boundary, because a third party's refusal is not a transient condition.
Interview question
Q: Your order saga has compensations for all five steps and they are all retried indefinitely. The business asks for a report of how much money is currently in the wrong place. What do you have to add, and what do you tell them about the sagas that are currently retrying?
What a strong answer covers: retry classification so permanent errors terminate · a terminal failed state with the amount at risk · the deadline set inside the external window · escalation to a worked queue with an age target · the honest answer that today the number is unknowable because "retrying" and "unrecoverable" are the same state, and the first version of the report is a scan of sagas retrying longer than their business window.
Quick check
Quiz: A refund compensation has been retrying for three days against a card scheme whose refund window is 48 hours. What is the design defect? Retries are not classified, so a permanent rejection is treated as transient and the saga never reaches a terminal state or a human.
Flashcard: What state is missing from a saga that retries compensation forever? — A terminal
compensation_failed state with the amount at risk, reached by a deadline set inside the external window, and
escalated to a queue with an age target rather than to a page.