pattern

Unmerge Capability

also called Merge Reversal, Split-Back

The designed ability to reverse an incorrect entity merge and restore both original records, which requires that source keys and pre-merge attribute values were never destroyed.

mdmentity-resolutiongolden-recordprivacylineage

Two customers share a surname, a postcode and a date of birth one of them entered wrongly. The matching engine scores the pair above the automatic-merge threshold and consolidates them. One person's order history and support tickets are now visible to the other, and an agent reading the golden record has no indication that two people are in front of them.

Whether this is recoverable was decided long before the merge, by whether consolidation was implemented as a reversible assertion over retained source records or as a rewrite that overwrote them.

Why it matters

A false merge is a privacy incident, not an inefficiency. A false non-merge leaves two records for one person and costs a duplicate mailing. The errors are not symmetric, which is why thresholds should prefer leaving pairs in the review band.

Given the asymmetry, the second line of defence matters as much as the threshold. False merges are certain at scale: a programme consolidating a few million records at a 99.9% precision target still creates thousands of wrong merges. A programme without an unmerge path has a known defect rate and no remediation, which is the finding that ends badly in an audit. Reversibility also buys throughput, because teams that can undo a merge accept a more aggressive threshold and clear the review queue faster.

Implementation patterns

  • Never overwrite a source record. The golden record is a derived view over retained source rows with survivorship applied; consolidation that mutates or deletes sources is irreversible by construction.
  • Keep the source key and system of origin on every contributing row, so each contribution's pre-merge identity is recoverable without reconstruction.
  • Record the merge as an event with the score, the matching configuration version and the responsible operator or job, so reversal replays it rather than guessing.
  • Issue new identifiers on an unmerge; never reuse the old ones. A reused identifier makes historical facts join to the wrong entity with no error anywhere.
  • Publish a mapping from superseded to surviving identifier with effective dates, so consumers resolve through it at a date they control rather than being renumbered under them.

Industry example

Insurance and banking consolidation programmes have treated reversibility as a hard requirement since regulators began asking how a customer's records are corrected on request. The European data protection regime that took effect in 2018 gives individuals a right to have inaccurate personal data rectified, and an estate that cannot reverse a merge cannot answer that request for the commonest inaccuracy it produces itself. The first wrong merge after a cross-brand consolidation reaches a complaints team within weeks.

Failure scenarios

  • Survivorship applied destructively. The losing record's address is discarded rather than retained as a non-surviving contribution, so unmerge cannot restore it.
  • Identifier reuse after a split, so every historical fact referencing the old identifier is ambiguous.
  • Downstream renumbering. The master is corrected and 200000 foreign keys in warehouse fact tables point at a superseded identifier, so joins drop rows or fan out and aggregates change with no pipeline failure.
  • Unmerge in the master only. Caches, search indexes and marketing platforms still hold the merged entity, so the person still sees the other's data in the channel that complained.

Trade-offs

Choose Gains Pays
View over retained sources Fully reversible; provenance is free Read-time survivorship cost; larger store
Materialised with provenance Fast reads and reversibility Storage for every non-surviving contribution
Destructive consolidation Smallest store, simplest reads False merges are permanent; restore from backup

When not to use it

When no record describes a person or an accountable party, reversibility is overhead. Consolidating a product catalogue or a list of warehouse locations has no privacy failure mode: a wrong merge is noticed by whoever looks at the product page, and re-running the match with a corrected rule is cheap.

It is also not worth building ahead of a matching engine. Where consolidation is exact match on a shared national identifier with no fuzzy scoring, the false-merge rate comes from source data-entry errors and case-by-case correction is cheaper. The capability becomes necessary once a probabilistic threshold decides merges without a human.

Interview question

Q: Your MDM platform merged two real customers at a score of 0.94 and a complaint has arrived. Walk me through the remediation, then tell me what in the original design determines how long it takes.

What a strong answer covers: confirming the false merge from retained source contributions rather than from the golden record; reversing by replaying the merge event, issuing two new identifiers and never reusing the superseded one; publishing the mapping with effective dates so downstream fact tables resolve rather than being renumbered; propagating the split to indexes, caches and downstream platforms, where the exposure actually lives; and that the determining choice was whether survivorship was destructive.

Quick check

Quiz: After unmerging two customers, why issue two new identifiers rather than reusing the original pair? Because a reused identifier makes historical facts join to the wrong entity with no error anywhere.

Flashcard: What single design choice decides whether a false merge is recoverable? — Whether the golden record is derived from retained source rows with full provenance, or was written by overwriting the sources.