intermediate 3 min answer

Customer phone numbers are replaced with vault-issued tokens at ingest, and the four services allowed to see raw values call the vault to detokenise on demand. The vault becomes slow rather than down - p99 goes from 8 ms to 4 seconds - for three hours during the evening peak. What happens across the estate, and what stops it?

tokenisationvaultavailability-couplingtimeoutsdegradation
Show the full answer Hide the answer

Second by second, what happens

Seconds 0 to 30. Nothing visible. The detokenise calls that used to take 8 ms now take 4 s, and the caller's thread pool absorbs the first few.

Seconds 30 to 120. The support console's connection pool saturates. A 50-thread pool at 4 s per call serves about 12 detokenise operations a second where it previously served well over a thousand — a 500× latency change converts directly into a 100× throughput collapse on that path, because the pool size was chosen for 8 ms work. Agents see spinners, press refresh, and each refresh issues a fresh request behind the ones already queued.

Minutes 2 to 10. Ingest is next, and this is the part designs usually miss. If tokenisation happens inline on the write path, the vault's latency is now the ingest pipeline's latency. The event queue grows at the arrival rate minus the degraded service rate, so a 10× evening peak against a 100× throughput loss means the backlog grows faster than it can ever be drained inside the three hours.

Minutes 10 onward. Retries amplify. Any client retrying a 4-second timeout without a budget doubles or triples the offered load on an already-saturated vault, which pushes p99 higher.

What keeps working, and why that is the point

The analytical estate is unaffected. Warehouse tables, dashboards, models and exports hold tokens and never call the vault, so a vault incident does not touch analytics at all. That asymmetry is the payoff of tokenisation and it is worth saying out loud in a review: the blast radius is exactly the set of paths that need raw values.

What stops it

  1. Bounded concurrency with a deadline shorter than the caller's budget. A semaphore of 20 in-flight detokenise calls and a 150 ms timeout means the console degrades to "number hidden - retry" instead of hanging. Slow becomes an error, which is what the caller can handle.
  2. Take the vault off the write path. Issue tokens from a deterministic, key-derived scheme with the key held in a hardware module, and record the reverse mapping asynchronously. Ingest then needs no round trip and the vault is a read-side dependency only. The cost is that deterministic tokens preserve joins and therefore leak equality, so the derivation input must have enough entropy to resist being rebuilt offline.
  3. A short-TTL local cache on the read side. A support agent looks at the same customer several times in a session, so a 60-second cache removes most repeat calls. It also widens the window in which a revoked permission still returns a value, so keep the TTL small and never cache in a shared tier.
  4. Shed at the vault, not at the caller. A bounded queue that rejects when full keeps p99 near normal for the requests it accepts. An unbounded queue converts overload into latency, which is what produced this incident.

When this is the wrong answer

When the only consumer of raw values is a nightly batch, remove the online detokenise path entirely. No cache, no semaphore, no capacity planning: a batch job reads the mapping once, and a three-hour vault incident costs a late report. Teams build highly available detokenisation because it feels like infrastructure, when the honest requirement was often "operations needs the number during a call" for a few dozen lookups an hour.

The self-healing condition is worth stating too. The system recovers on its own only if every caller fails fast and no caller retries without a budget. With unbounded queues anywhere in the chain, the backlog outlives the trigger and someone has to drain it by hand.