Resync Horizon
also called Tombstone Retention Window, Full Resync Threshold
The age of the oldest change a server can still describe as a delta - beyond which a returning client must take a full snapshot or it will silently keep records that were deleted.
An engineer returns from four months of leave, opens the field app, and syncs. The request succeeds, the client's cursor is accepted, and jobs cancelled in February reappear on the device — then, on the next upload, in the shared database for everyone. No error is logged on either side.
The cause is arithmetic rather than a bug. Deletions are communicated as tombstones in a change feed, the feed is pruned to stay fast, and a cursor older than the pruning point cannot be answered correctly. The horizon is the boundary between a cursor that can be served and one that can only be refused.
Why it matters
Delta sync is the standard way to make a mobile client cheap: send a cursor, receive what changed. The protocol's correctness rests on a premise nobody writes down — that the server still remembers everything that happened since that cursor — and the premise is quietly broken by a retention job added later for performance.
The failure is silent in both directions and it is worse than data loss, because the resurrected records look authentic. They have real identifiers and real history, they flow into reports and billing, and the only people who can tell they should not exist are the users who cancelled them months ago.
Implementation patterns
- Store the horizon as a value, not a convention. The server keeps a watermark for the oldest change it can describe, and the pruning job advances it in the same transaction that prunes.
- Refuse, do not guess. A cursor older than the watermark gets an explicit
resync_requiredresponse. The client drops local state for that scope and takes a snapshot. - Scope it. Horizons are per tenant or per collection, because retention and change volume differ; one global number is either wasteful or wrong.
- Size retention from the gap distribution, specifically the p99.9 of device reconnect gaps. Seasonal staff, shared handsets and per-vehicle devices produce six- to nine-month tails.
- Prefer keeping tombstones. At roughly 40 bytes each, 2 million deletions is about 80 MB. Most pruning jobs are saving storage that was never worth saving; index the feed instead.
- Never treat an unknown key on upload as a create. A genuine creation carries a creation marker; anything else is an update that may legitimately fail.
- Make snapshot resync cheap and resumable, because it is now a normal path rather than a disaster-recovery one.
Industry example
Retail and grocery fulfilment platforms, of the kind an operation in Instacart's mould runs, synchronise large catalogue and inventory sets to shopper devices where items are delisted constantly. The characteristic production bug is not conflict resolution but exactly this: a handset left in a locker for a quarter, returning with a cursor the server can no longer honour, offering back items the retailer removed. The pattern generalises to any change-data-capture consumer — a stream with a retention period of seven days has a horizon of seven days, and a consumer that lags past it must be rebuilt from a snapshot rather than resumed.
Failure scenarios
- Resurrected deletions, as above, re-uploaded by the client as creates.
- A permission revocation that never lands, because the revocation was a deletion in the feed and the device missed it, leaving data visible on a handset that should no longer show it.
- A retention job added for performance with no awareness that it is a protocol change.
- Cursor age never measured, so the first evidence is a customer report and the correlation (long absences) sits in a system the engineering team does not look at.
- Resync implemented but not resumable, so the device that most needs it is the one that cannot complete it on a weak link.
- A horizon shorter than the app-store update tail, where an old client version cannot even understand the resync response.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Long tombstone retention | Any plausible absence is servable as a delta | Storage and a feed that grows, needing indexing and partitioning |
| Short retention plus enforced resync | A small fast change feed | A snapshot path that must be cheap and resumable for the worst-connected devices |
Enforcement has a user-visible cost: a returning device spends minutes and megabytes rebuilding its state, on exactly the connection least able to afford it. That is the honest price of correctness, and it is why retention long enough to make resync rare is usually the better engineering choice even when it looks less elegant.
When not to use it
If every client synchronises at least weekly and the feed is retained for a year, the resync protocol is complexity without a case; keep the tombstones and skip it. The horizon matters when the gap distribution has a tail you cannot shorten — devices that live in vehicles, lockers or the hands of seasonal staff. For a server-rendered product with no local store, there is no cursor and no horizon.
Interview question
Q: Your change feed is pruned after 30 days and a device returns after eight months. Describe exactly what the server should reply, what the client must do, and the alert that would have warned you a month before a customer noticed.
What a strong answer covers: the watermark and an explicit resync response rather than an honest answer to an unanswerable cursor; the client dropping scoped local state and taking a resumable snapshot; the upload rule that an unknown key is not a create; sizing retention from the p99.9 reconnect gap; and the alert, which is a histogram of accepted cursor ages firing at 80% of the horizon.
Quick check
Quiz: Why does answering a too-old cursor literally cause data corruption rather than staleness? Answer: because the deletions in that window have been pruned, so the client keeps deleted rows and, on its next upload, reinserts them for every other client.
Flashcard: A device returns with a cursor older than your tombstone retention. What does the server owe it? — A resync_required response and a resumable snapshot; answering the delta honestly resurrects deleted records fleet-wide.