advanced 3 min answer

Eighteen months into a coexistence period, customer records are authoritative in the new system and order records are still authoritative in the old one. A refund credits the customer and cancels the order line. What happens the first time the credit commits and the cancellation is rejected?

coexistencewrite ownershipsagascompensationreconciliation
Show the full answer Hide the answer

Second by second

The refund handler writes the credit to the new system, which commits. It then calls the old system's order API to cancel the line, and that call returns a rejection, because the order shipped four minutes earlier and the legacy validation forbids cancelling a dispatched line. There is no transaction spanning the two systems, so nothing rolls the credit back. The handler logs the rejection and returns an error to the operator, who retries. The retry finds the credit already present only if the handler was written with an idempotency key; if it was not, the second attempt credits the customer again.

What was a single ACID unit inside the monolith is now two writes with a network between them, and the ownership split created that seam without anyone deciding to.

Where it amplifies

Three multipliers, in the order they bite:

  1. Retry. Refunds usually run from a queue with automatic retry. A rejection that is permanent rather than transient turns into repeated compensating writes until the dead-letter threshold, so one bad case becomes three to five duplicate credits.
  2. Volume. Cross-boundary business transactions are a minority of traffic but not a rare one. At roughly 1% of orders needing a refund and a few percent of those racing a dispatch, a platform doing 50,000 orders a day produces a handful of divergences a day — enough to be a permanent finance ticket queue and too few to trip an error-rate alert.
  3. Time. The coexistence period is eighteen months, so the divergence accumulates for as long as the programme runs.

What the user sees

Nothing. That is the problem. The customer sees a credit. The warehouse sees a shipped order. The first place the two facts meet is the monthly finance reconciliation, weeks later, in aggregate, where the error appears as an unexplained balance rather than as a list of broken transactions.

What stops it

Draw the write-ownership boundary along transaction boundaries, not along entity names. If a single business transaction touches customer and order, one system owns the whole transaction for the duration of coexistence, even if that means the new system stays a read-only view of orders longer than the roadmap wanted. Ownership drawn on a data model diagram will cut straight through transactions, because the data model was never drawn to show them.

Where a split is unavoidable, make it explicit rather than accidental:

  • The cross-boundary flow becomes a saga with a named compensating action for each step, not a sequence of calls with an error log.
  • Every step carries an idempotency key derived from the business event, so retries are safe.
  • The compensation failure path has its own dead-letter queue that pages, because a failed compensation is a correctness defect and not a backlog item.
  • A daily reconciliation between the two systems on the entities that span the boundary, with the count of unmatched pairs as an alert on the first occurrence.

For this to self-heal without any of that, the old system would have to expose a reservation the new system could hold and release, which is exactly the two-phase behaviour most legacy APIs do not offer.

When this is the wrong answer

If the two entities are genuinely independent, split ownership by entity and move on. Customer marketing preferences and order history do not participate in a common transaction, and adding a saga there is machinery protecting nothing. The rule binds only where a single business transaction spans the boundary, and the way to find those is to list the operations that used to run in one database transaction, not to inspect the schema.