Replica Id Inflation
also called Actor Id Growth, Unbounded Causal Context
The steady growth of coordination-free merge metadata driven by the number of distinct replica identifiers a record has ever seen - which makes a document slower every year even when editing volume is flat.
Two years after launch, heavily used records in a collaborative field app take 40 seconds to open and the local database is hundreds of megabytes. Editing volume has not changed. The growth is not content and it is not tombstones: it is the causal context, and its size is driven by how many distinct replica identifiers have ever touched the record.
Every coordination-free merge needs to know which updates it has already seen, and it expresses that as a map from replica identifier to counter - a version vector, or per-element dots. The map cannot shrink on its own, because the algorithm has no way to distinguish a replica that was uninstalled last year from one that has been offline for a year and may still arrive carrying concurrent updates. Dropping an entry risks resurrecting deleted data or losing a live edit, so correct implementations keep it.
Then the identifier churns. A fresh install, a cleared data directory, a factory reset, a restored backup and a second device all create new replica identifiers for the same human. A user who reinstalls four times a year contributes four permanent entries. At roughly 16 bytes of identifier plus a counter, a record touched by 300 historical replicas carries about 6 KB of vector before any content, and a type that keeps per-element dots multiplies that by element count.
Why it matters
The symptom appears long after the design decision and looks like nothing in particular. No error, no conflict, no failed sync - just a product that is slower every quarter, worst for the most engaged users, whose records have the most history. Teams chase query plans and storage engines for months before measuring metadata as a share of bytes.
It also sets a ceiling on what the architecture can promise. "Works offline forever" and "bounded local storage" are in direct tension once causal metadata is unbounded, and the tension is resolved by a product decision about horizons, not by a library.
Implementation patterns
- Server-assigned stable replica ids bound to account plus device rather than to the install, so a reinstall reuses its identifier. This single change removes most churn in consumer apps.
- Server-side compaction with a rebase. Once every live replica has acknowledged a state, write a snapshot with a fresh causal context and have replicas adopt it. This is coordination, deliberately reintroduced for garbage collection, and it is why a server-authoritative model is simpler when a server already exists.
- A published replica retirement horizon - for example 90 days - after which a returning replica must take a snapshot rather than merge. Say the number in the design, because it is the same horizon a resync protocol needs.
- Causal stability tracking. Compute the minimum across all known replicas' vectors and discard metadata below it, which requires knowing the replica set and therefore a membership protocol.
- Measure metadata share directly. Report bytes of causal context versus bytes of content per record, per cohort. If metadata exceeds roughly 30% of a record, the design has a year or two before it is the product's main performance problem.
Industry example
The literature says this plainly. The original CRDT paper (Shapiro and colleagues, 2011) establishes convergence; the work that follows it - delta-state CRDTs (2016) and causal stability - exists because bounded metadata is the hard part, not an implementation detail. Convergence is the easy property; garbage collection is the one that needs a protocol.
It is also why commercial collaborative editors that have published their multiplayer designs tend to describe a server-authoritative model inspired by CRDTs rather than a pure peer-to-peer one. A central server removes the need for coordination-free garbage collection, because it can decide what history may be dropped and when every live replica has caught up. If a server is already in the request path, taking that help is the cheaper engineering decision.
Failure scenarios
- Open latency grows with tenure. The oldest, most valuable accounts are the slowest, and the correlation is invisible in aggregate percentiles.
- Sync payloads grow faster than content. A reconnect ships a vector larger than the changes it describes, which on cellular is a visible data-plan line.
- A reinstall "fixes" the user and worsens the record. Support advises reinstalling, which adds another permanent identifier to every record that user touches.
- Compaction that is not safe. A team prunes the vector to control growth, and a device that returns after three months resurrects deleted items or has its edits dropped silently - the worst class of bug in a sync system, because users cannot report what disappeared.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Pure coordination-free merge | Offline writes with no server in the merge path | Unbounded causal metadata and no safe garbage collection |
| Server-mediated compaction and rebase | Bounded metadata and a horizon you control | A server in the path and a membership protocol |
| Last-writer-wins registers for most fields | Near-zero metadata | Concurrent edits to the same field are discarded without trace |
When not to use it
If replicas are a small, stable set of servers rather than a churning population of phones, the vector is bounded and none of this machinery is needed. Likewise if documents are short-lived - a session, a shift, an order - the metadata never has time to accumulate, and adding compaction is cost without benefit. And where one writer owns a field, a last-writer-wins register with a server-assigned sequence is the right answer; the cheapest way to avoid replica id inflation is to use a coordination-free type only where concurrent edits are genuinely information.
Interview question
Q: A CRDT-backed app's records grow every quarter while editing volume is flat, and the worst records belong to users who reinstall often. Explain the mechanism and describe a fix you would be willing to ship.
What a strong answer covers: causal context sized by distinct replica ids; why entries cannot be dropped without a membership or stability protocol; install-churn as the source of new ids; stable server-assigned replica ids as the low-risk first fix; snapshot-and-rebase compaction once all live replicas have acknowledged; a published retirement horizon tied to the resync horizon; metadata-share measurement as the leading indicator; and the honest observation that a server in the path makes all of this easier, which is why server-authoritative designs are common where a server exists.
Quick check
Quiz: Why can a version vector entry for a replica that has not been seen in a year not simply be deleted? Because the algorithm cannot tell a retired replica from an offline one, and a returning replica's updates would then merge incorrectly - resurrecting deletions or dropping edits.
Flashcard: What drives CRDT metadata growth when editing volume is flat? The number of distinct replica identifiers the record has ever seen - inflated by reinstalls, resets and restored backups - at roughly 16 bytes plus a counter each, kept forever unless a stability or membership protocol allows compaction.