Boundary Reversal Asymmetry
also called Split Reversal Cost, Asymmetric Boundary Error
The fact that a boundary drawn too coarse is cheap to split later while one drawn too fine is expensive to merge, which makes the coarse error the right default when evidence is thin.
A team splits checkout from pricing because the two sounded like different concerns. Six months on, every checkout request makes three calls into pricing and a promotion change needs both services released in order. Merging them back means deleting a public interface four other services now call, and folding two schemas into one.
Another team left ordering and fulfilment together, read version-control history, saw the halves changing in different weeks for different reasons, and split them in a month behind a facade.
Both teams drew a boundary without enough evidence. Their recovery costs differ by roughly an order of magnitude, and the difference runs in one direction.
Why it matters
The evidence that settles a boundary is change history and load, and both exist only after the system has run. The practical question is therefore not where the boundary belongs but which error to prefer while you find out.
A too-coarse boundary keeps its cost inside one codebase. Data, transaction and tests sit in one place, so splitting is a refactor you can sequence behind an interface and abandon half way.
A too-fine boundary spreads its cost into things no single team can refactor. A transaction becomes a saga and a reconciliation job. A method call becomes an in-datacentre round trip of roughly 0.5 to 1 ms instead of a fraction of a microsecond, with a timeout and a retry policy attached. One pipeline becomes two, and the interface becomes other teams' dependency. That last one is the expensive item: you cannot delete an interface four teams call.
Implementation patterns
- Default to one service per team, with enforced modules inside. The module boundary is the hypothesis; the service boundary is the commitment.
- Promote a module to a service only on evidence: halves changing on different cadences, a load difference big enough to want separate capacity, or a compliance scope to shrink.
- Keep the first split's interface private to one consumer for its first quarter. One caller to delete is a reversible split; six callers is not.
- Never split an invariant that must hold exactly. The split converts a database guarantee into an application problem you will be reconciling for years.
- Write the reversal plan at the split: what gets deleted if this is wrong, and who decides by when.
Industry example
Martin Fowler's "MonolithFirst" (June 2015) is the empirical version of the argument: nearly every successful microservice system he could find began as a monolith that was broken up, while systems built as fine-grained services from the start repeatedly ran into serious trouble. Starting coarse leaves the decision open until the system answers it; starting fine commits to a guess and makes the guess expensive to revisit.
The archetypal expensive version puts the payment charge and the ledger entry on opposite sides of a boundary. Nothing fails loudly. The symptom is a nightly reconciliation report with a mismatch count nobody can explain, and an engineer owning the repair job.
Failure scenarios
- Lock-step releases. Two services that cannot deploy independently paid for distribution and bought nothing.
- The interface that outlives the mistake. Consumers appear faster than you can retract the split, and the split invariant shows up as a mismatch count rather than an error.
- The organisational lock. Once two teams own the halves, merging means changing the org chart.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Coarse first | Cheap splits later, one transaction, one deploy | Contention between teams in one codebase, a bigger blast radius per release, enforcement work to keep modules honest |
| Fine first | Independent scaling and ownership from day one | Expensive merges, sagas, two of every operational thing, a guess baked into other teams' code |
When not to use it
The coarse default is wrong when the two sides differ on something a shared deployable cannot express. A workload needing specialised hardware, code you run on behalf of untrusted users, a component you want outside a regulated audit scope, or a hundredfold load difference all justify the fine cut up front, because the coarse version forces the strictest and most expensive regime on everything in the process.
It also flips when the boundary is already an organisational fact: if teams from two merged companies own the halves, coarse is not available at any price.
Interview question
Q: You must choose a decomposition now with no production history. Argue for the coarser option, then tell me what evidence would make you split, and what you would put in place today to make that split cheap.
What a strong answer covers: the asymmetry in recovery cost; the evidence that resolves the question later (co-change history, divergent load, compliance scope, team ownership); the preparation that makes a later split cheap (enforced module boundaries, a schema per module, no cross-module joins); and the cases where the fine cut is right on day one.
Quick check
Quiz: Why is "too coarse" the safer error when boundary evidence is thin? — Splitting later is a refactor inside one codebase; merging later means deleting an interface other teams depend on, roughly ten times the work.
Flashcard: Why is a wrong fine-grained boundary more expensive than a wrong coarse one? — Its cost sits in what you cannot refactor alone: other teams' clients, a split invariant and two operational surfaces.