Token Domain Entropy
also called Tokenised Value Space, Surrogate Input Entropy
The size and guessability of the value space a deterministic token is computed from, which decides whether that token is a usable pseudonym or a lookup key an attacker can rebuild without ever touching the vault.
Deterministic tokenisation is adopted for one reason: analysts still need to join. Replace every mobile number with a stable surrogate and the orders table still joins to the support table, while no raw number sits in the warehouse. The property that makes the join work, the same input always producing the same token, is also the property that can hand the whole mapping to an attacker.
Whether it does depends almost entirely on the input domain, not on the algorithm. A ten-digit mobile number with a constrained leading digit has a candidate space on the order of 1 billion to 5 billion and a realistically reachable space far smaller. A national identity number with a checksum and an embedded date of birth is smaller still. A 128-bit random account reference is not enumerable by anybody.
Token domain entropy names that property so it can be argued about in a design review, where the conversation otherwise stalls on whether the algorithm is strong. The algorithm is rarely the weak part.
Why it matters
The attack needs two things and neither is the vault. It needs a low-entropy input domain, and it needs any path that will tokenise a value the attacker chooses: a signup flow that echoes a token back, a support form, a batch ingestion endpoint, an internal tokenise service reachable from the analytics network without strong authentication. Given both, building the dictionary is a loop, and the vault's audit log stays empty throughout.
The blast radius is then the entire retained history, because the token is stable across every table and every snapshot. That is not an implementation slip; it is the feature that was purchased.
It also settles a compliance question that teams get wrong. Pseudonymisation is not anonymisation. GDPR Article 4(5) defines pseudonymisation as processing after which the data can still be attributed to a person using separately held information, and Recital 26 keeps such data within scope as personal data. In the 2016 text applicable from 2018, a tokenised warehouse is a warehouse of personal data with better controls, not a warehouse outside the regime.
Implementation patterns
- Classify the input domain before choosing the scheme. Estimate the reachable candidate count. Anything under roughly 2^40 that an attacker can also enumerate should be treated as re-identifiable.
- Treat tokenise as a privileged operation: authenticated, rate-limited per caller, and monitored on volume and caller distribution. A chosen-input oracle is the precondition for the whole attack and it is the cheapest thing to remove.
- Scope determinism to the join that actually exists. A per-dataset key gives joins inside a dataset and no cross-dataset dictionary. Most analytical joins turn out to live inside one domain.
- Prefer a surrogate identifier assigned at first sight for the analytical estate, with no functional relationship to any attribute. There is nothing to enumerate because the mapping is not computed from anything.
- Keep format-preserving schemes for operational compatibility only, and accept that preserving format usually means preserving domain size.
- Log detokenisation separately from tokenisation and alert on either in the wrong direction.
Industry example
Characteristic of payments platforms of the kind Razorpay operates rather than attributed to any: a gateway tokenises customer mobile numbers before analytics, deterministically, so support tickets join to orders. Nine months later a red team rebuilds the mapping for a large share of the customer base using an internal tokenise endpoint that required no authentication from inside the network. The vault was never touched, no access review flagged anything, and the only telemetry that would have shown it was the request rate on a service classified as a data-preparation utility.
Failure scenarios
- Salted hashing used as a pseudonym. The salt is not secret from anyone who can compute hashes against a candidate list, so it changes the dictionary's cost by a constant and nothing else.
- Key rotation as the incident response. It breaks every historical join and every model keyed on the token, and it does not prevent the same attack against the new key.
- The token becomes a business key. Once partners and downstream systems key on it, the token is as hard to change as the raw value was, and the compliance boundary has moved rather than shrunk.
- Tokenising the identifier while leaving a quasi-identifier set intact. Postcode, birth date and gender in the same row re-identify a large share of a population without touching the token at all.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Deterministic token on the real identifier | Joins work everywhere with no data model change | Re-identifiable whenever the input domain is enumerable and an oracle exists |
| Surrogate id assigned at first sight | Nothing to enumerate; mapping stays in the operational system | A resolution step on ingestion and a model change for every consumer that joined on the old value |
When not to use it
The concern does not apply when the input is genuinely high-entropy and not chooseable: a long random account reference, an internally generated opaque id. There, deterministic tokenisation is appropriate and the surrogate rewrite is wasted effort. It is also the wrong lens for free-text fields, where the risk is content rather than key space and the answer is redaction or exclusion, not a better token.
Interview question
Q: Your warehouse holds deterministically tokenised phone numbers so that analysts can join across datasets. Argue whether that is adequate protection, and say what evidence would change your answer.
What a strong answer covers: that the algorithm is not the question and the input domain is; estimating the reachable candidate space; identifying whether any path tokenises attacker-chosen input; the permanence of the exposure given a stable token across history; why rotation is the wrong response; the surrogate-id alternative and what it costs consumers; and the legal position that pseudonymised data stays personal data under GDPR Article 4(5) and Recital 26.
Quick check
Quiz: Why is a salted hash of a phone number not a safe pseudonym? The salt is not secret from anyone who can compute hashes over the ten-digit candidate space, so it multiplies the attacker's cost by a constant while the domain stays enumerable.
Flashcard: Which two conditions together let someone rebuild a deterministic token mapping without touching the vault? An enumerable input domain, and any path that will tokenise a value the attacker chooses. Remove either one and the dictionary attack stops being feasible.