Review this. A grocery-fulfilment platform in the mould of Instacart runs an orchestrated saga over five steps - cart freeze, pricing, promotion, loyalty accrual, order creation - each with a compensating action, driven by an orchestrator with its own saga log. All five services are owned by one team, deployed as one binary, and write to five schemas in one Postgres cluster. What would you remove, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
A saga exists because no transaction can span the participants. Exactly three conditions create that, and each is checkable in a few minutes:
- Separate transactional stores, so there is no commit that covers both.
- A party you do not control, where you cannot hold a transaction open across the boundary.
- A step with an external effect, which cannot be rolled back because something outside the system already observed it.
Run the five steps against the three conditions. One Postgres cluster, so not the first. One team and one binary, so not the second. Cart freeze, pricing, promotion and loyalty accrual all write rows the platform owns, so not the third. The table is empty, which means the saga is machinery that buys nothing.
What I would remove, and why it is safe
The orchestrator, the five compensating actions, the saga log and the five broker round trips. Five schemas in one Postgres cluster participate in one transaction without ceremony, so the forward path becomes a single BEGIN and COMMIT.
The saving is easy to state. Code paths drop from ten - five forward, five compensating - plus every partial-failure interleaving the test suite ought to cover, to one. The broker hops and the orchestrator's own state writes are on the order of 50 to 150 ms of p50 that disappears. And the system gains what the saga never provided: isolation. A saga gives atomicity and durability and no isolation, so another request can read a half-applied order; the database gives all three.
The one change that matters
Payment. If the capture call goes to an external provider, that single step satisfies condition three and is the only real boundary in the design.
Keep a two-step saga there, and only there. The local transaction commits the order in pending and writes an outbox row. A relay drives the capture with an idempotency key derived from the order id. The compensating action is a refund, which is a new business action rather than a rollback, and it needs an owner because a failed compensation is a correctness defect with no automatic recovery.
What I would leave, even though it looks odd
- The saga log entry for the payment step. It looks inconsistent to keep a saga log for one step, and it is the only durable record of an external effect whose outcome may arrive minutes later or by webhook.
- The idempotency key, which is doing work the single transaction cannot: the provider may have captured before the timeout you saw.
- The five schemas. Collapsing them to "simplify" is the one piece of this cleanup that costs money later. Schema separation inside one cluster is nearly free and is the thing that makes extracting a service a deployment change rather than a data migration. Keep the boundary, delete the distribution.
How I would argue this in the review
Not as "sagas are over-engineering", which reads as taste. Put the three boundary conditions in a table against the five steps and let the empty cells make the case. Then price it in what the team feels: the pages for this service come from compensation failures and stuck saga instances, and both categories go to zero.
When not to collapse it
If any one of the three conditions holds for any step, the saga stays. A team split that moves those schemas to separate clusters brings it back, as does a step that sends a customer email, because an email cannot be rolled back.
So make the collapse reversible on purpose: put the single transaction behind one application-level service so that reintroducing coordination is a change in one place rather than five. The failure mode of this review is a team that collapses the saga, splits six months later, and reintroduces distributed coordination by accident in five call sites.