advanced 4 min answer

Review this: a bank generates synthetic test data by taking a production dump and replacing names, emails, account numbers and dates of birth with random values. The data volume matches production, the schema is identical, and the compliance team has signed it off as anonymised. What would you change?

synthetic dataanonymisationre-identificationreferential integritytest data
Show the full answer Hide the answer

What is actually required

Two things, and the design in the stem conflates them: test data that exercises the system realistically, and data that cannot be traced back to a person. Field-level masking of a production dump is a weak attempt at the second and quietly damages the first.

The one change that matters

The sign-off is wrong, and that is the finding to lead with. Replacing direct identifiers is pseudonymisation, not anonymisation, and the distinction is legal as well as technical. The dataset retains every quasi-identifier: postcode, transaction amounts, timestamps, merchant names, branch, account-opening date, balance trajectory. A transaction history is close to unique per person. Anyone who knows three transactions a person made — which a merchant, a colleague or a family member does — can locate that person's row and read everything else, including the attributes that were never masked because they were not considered identifying.

Under the UK and EU regimes, pseudonymised data remains personal data and carries the full set of obligations. A sign-off calling it anonymised is a control failure, not a paperwork detail: it will have been relied on to put this data somewhere it does not belong, and that is the first thing to find out.

What I would change, in order

  1. Re-scope the sign-off and establish where the dataset currently lives. Non-production environments typically have weaker access controls, longer retention and more copies than production, so the exposure is usually wider than anyone expects.
  2. Preserve referential integrity deliberately, or the data does not test anything. Randomising per-field almost certainly broke the joins: the same customer now has different identifiers in different tables, so any query spanning them returns nothing and the tests that pass over empty result sets pass vacuously. A test suite that is green against broken joins is the worst outcome available here — slower to detect than a red one and more confidently believed. Masking must be deterministic and consistent: the same input maps to the same output everywhere.
  3. Preserve distributions, not just types. Random values destroy the properties that make the data a useful test: the skew where 2% of accounts hold 60% of transactions, the seasonal pattern, the correlation between balance and product holding. Query plans, cache hit rates and partition hotspots all depend on distribution, and uniform random data makes every one of them unrepresentative.
  4. Preserve the awkward edge cases. Real data contains the account opened in 1987 with a null field that validation now forbids, the name with an apostrophe, the four-byte emoji in an address line, the negative balance. These cause a large share of escapes and they are precisely what random generation omits.
  5. Decide between two honest strategies rather than the hybrid in the stem, which has the drawbacks of both: - Generated-from-model data — learn the schema, the distributions and the constraints, then synthesise fresh records. No individual is present, so the re-identification question does not arise. Costs real engineering effort and never quite captures the oddities. - Properly de-identified production data — formal treatment of quasi-identifiers through generalisation, suppression or a differential-privacy mechanism, with a documented re-identification risk assessment. Keeps more realism and demands expertise and ongoing review.
  6. Shrink the volume for most purposes. Production-scale data is needed to test query plans and capacity, and almost nothing else. A 1% representative extract serves the functional suite, costs a fraction, and reduces the exposure proportionally. Keep the full-size set for the performance environment alone.

What I would leave, even though it looks odd

Keeping production-matched data volume for the performance environment specifically. It looks like gratuitous risk and it is the one place scale is irreplaceable: query plans change discontinuously with table size, and no amount of distributional fidelity in a small dataset reproduces that. Keep it, restrict access to a named list, and justify it on exactly that basis rather than on general realism.

When not to rebuild the generator

If this dataset is used only by the performance environment, under a named access list, and the functional suites run on a small generated set already, then the re-identification exposure is bounded and the broken joins do not matter — nothing in a load test asserts on a join's contents. Fix the sign-off wording, restrict the access, and leave the data alone. The expensive recommendation above is justified by the data being used for functional testing, so establish that it is before proposing a quarter of generator work.

How I would argue this in the review

Lead with the legal exposure, because that is the argument that moves a bank. Then demonstrate the broken joins concretely: run one cross-table query against the masked set and show the empty result, which turns "the data is realistic" into a question with an answer in front of everyone. The two findings reinforce each other — the data is both riskier and less useful than believed — and that combination is what unlocks funding for the generator, which is otherwise a hard thing to get approved.