A delivery app syncs three things from each courier's phone. An availability toggle the courier flips; a cash-collected ledger of per-order amounts; and a per-order delivery status. Couriers go offline for up to an hour in basements and lifts, and a minority have the app installed on two phones. Which conflict policy fits these three fields?
Show the full answer Hide the answer
The deciding property
For each field, ask two questions: how many writers can legitimately produce a different value, and does a later write supersede an earlier one or add to it? The answers differ across these three fields, which is why one policy for the record is wrong.
- Availability toggle. One writer, the courier. A later flip genuinely replaces an earlier one. Nobody wants to merge "available" with "unavailable".
- Cash ledger. Each entry is a distinct fact about a distinct order. Nothing supersedes anything, and losing an entry loses money.
- Delivery status. One writer, but not all orderings are legal. Delivered must not be overwritten by in-transit even if in-transit was sent later by a phone whose clock is wrong or whose queue drained late.
Why this split
LWW for the toggle, ordered by a server-assigned sequence rather than a device timestamp. A device clock is user-editable input; a 40-second fast clock on a second phone makes a stale toggle win permanently. The server's receipt order is imperfect but it is monotonic and it is yours.
Append-only for the ledger, with a client-generated idempotency key persisted to disk before the write is attempted. The phone queues entries offline and replays them; the key makes replay harmless. The balance is derived by summing, never stored as a mutable total, so a duplicate submission is a no-op rather than a double-count.
A state machine for status. The server holds the legal transitions and rejects the rest, returning the authoritative state so the device can correct itself. Convergence is not validity: both LWW and a CRDT will happily converge on "in transit" for a parcel that was delivered 20 minutes ago, because neither knows that the transition is illegal. This is the field that shows why conflict resolution and domain rules are not the same problem.
Why the other options fail
- LWW on all three. It is right for the toggle and silently destroys the ledger: two orders collected while offline become one, and the loss is money with no error anywhere. It also permits the illegal status regression.
- One CRDT document. It converges and it is the wrong tool for two of three fields. A CRDT cannot enforce an invariant such as "delivered is terminal" or "the ledger must reconcile with the cash handed in", because enforcing an invariant across replicas needs the coordination a CRDT exists to avoid. You also take on metadata growth and a client library for a problem that a sequence number and an append-only table solve. Server authority is cheaper whenever a server is already in the path and one writer owns the field.
- Conflict copies everywhere. The only strategy that never loses work, and it puts a merge interface in front of a courier holding a parcel in the rain. Reserve it for fields where both sides are genuinely information, such as free-text notes.
- A distributed lock. It cannot be held across an hour offline, and if it could, the courier would be unable to work when the network is gone, which inverts the requirement.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| Two devices legitimately edit the same free-text note | A text CRDT or multi-value register for that field | Both edits are information |
| The ledger must support corrections and deletions | Append-only plus tombstones and a resync horizon | Deletes need their own convergence story |
| Offline windows grow to days | Leases on offline capability | Stale authorisation becomes the bigger risk than stale data |
| A field gains a second legitimate writer such as dispatch | Move it behind the state machine | Two writers plus LWW is where data quietly disappears |
When not to split by field
If every field in the record is written only by the server, there is no conflict policy to choose - the device holds a cache and the answer is a change feed plus a cursor. The per-field split is worth its complexity only where the device writes.
The other case to avoid it: a single-device, always-online product where offline windows are seconds rather than an hour. Then the simplest correct design is optimistic concurrency on a version column, with the server rejecting a stale update and the client refetching. One version check beats three policies, and it fails loudly, which is what you want while the product is still changing shape.