Pseudonymisation
Separating identity from record, and the re-identification risk that remains.
4 to work through
-
intermediate Multiple choice
A product analytics team needs per-user aggregation over two years of events without holding data that re-identifies people if the warehouse leaks. Which approach actually delivers that?
2 min answer -
advanced
A travel metasearch of Expedia's shape replaces every traveller email with a deterministic HMAC token before the data reaches the analytics warehouse and before any file goes to an advertising partner. Two years later an audit finds that one partner can name individual travellers. The token vault was never breached. What failed?
3 min answer -
advanced
On 4 August 2006 AOL published about 20 million search queries from more than 650,000 users over three months, with usernames replaced by random numbers. Within days the New York Times had identified user 4417749 as a named individual from the content of her searches alone. What failed, which design assumption made it possible, and what would have prevented it?
3 min answer -
advanced
Pseudonymisation reduces regulatory obligations. What does it actually provide, and where is it commonly overstated?
2 min answer
4 terms in this topic
Key-Custodian Separation
Keeping the mapping or key that re-identifies pseudonymised data in a system with a different access path from the data itself - so a breach of the a…
practicePseudonymisation
Replacing identifying fields with a reference so records cannot be attributed to a person without separately held additional information.
conceptQuasi-Identifier Set
The combination of attributes that singles out an individual even after direct identifiers are removed - which is why deleting names and ids does not…
conceptRe-Identification Risk
The probability that pseudonymised data can be linked back to individuals, which is what keeps such data within the scope of data protection law.
Neighbouring topics
Regulatory & Data Protection Architecture
General material on designing under legal and regulatory obligation.
Privacy by Design
Data minimisation, default protection and purpose binding as structural decisions.
Lawful Basis & Purpose Limitation
Why you may hold the data, and why that forbids the second use somebody proposed.
Data Subject Rights
Access, correction, portability and erasure across systems that never planned for them.
Consent Architecture
Capturing, versioning and propagating consent to every system that acts on the data.
Data Residency
Keeping data within a jurisdiction, including backups, logs and support access.
Digital Sovereignty
Control over data, operations and the operator, beyond where the bytes physically sit.
Cross-Border Transfer
The legal mechanism that permits data to leave, and the architecture that respects it.
PCI-DSS Scoping
Segmentation and tokenisation to shrink what is in scope, because scope is the cost.
Healthcare Data Protection
PHI handling, minimum necessary access, and audit expectations in clinical systems.
Financial Services Regulation
Operational resilience, payment rules and supervisory expectations as design inputs.
Records Retention & Legal Hold
Keeping what must be kept, deleting what must go, and freezing both on demand.
Erasure vs Immutability
Deletion obligations against event logs, backups and ledgers designed never to forget.
Privacy-Enhancing Technologies
Differential privacy, secure enclaves and federated computation, and what each buys.
Regulatory Reporting Pipelines
Submissions with fixed deadlines, fixed formats, and a regulator who audits the lineage.
Sector Cloud Rules
Regulator expectations for cloud use, exit plans and material outsourcing notification.
Exit & Concentration Risk
Being able to leave a provider, and what the regulator asks when you cannot.
Third-Party Risk
Assessing, contracting and monitoring the vendors your architecture now depends on.
Geo-Restriction & Sanctions
Blocking access by jurisdiction, and the accuracy and evasion problems that come with it.