A marketplace of eBay's shape has an append-only event log of roughly 8 billion events over six years with personal data inline in the payloads. A new erasure obligation lands and crypto-shredding is the chosen answer, but no per-subject keys exist yet. Sequence the migration under live traffic and name the point of no return.
Show the full answer Hide the answer
The sequence
- Key registry before anything writes. A per-subject data key in a managed key store, wrapped by a per-region master key, with the key id derivable from the subject id so no lookup table becomes the single point of failure. Issue lazily on first write. This step has no observable effect and is fully reversible.
- Readers first, writers second. Deploy consumers that accept both the plain and the enveloped payload
shape, and only then switch producers to write
{key_id, ciphertext}for the personal fields. Leave non-personal fields in the clear so projections that never touch identity keep working unchanged. Reversing the order strands consumers on a format they cannot parse. - Backfill by rewrite into a new stream, not in place. An append-only log is append-only, so history
cannot be re-encrypted where it sits. You either write a parallel stream with a new offset space plus an
old_offset → new_offsetmapping, or you accept that pre-cutover history stays readable until its retention expires. Most teams should choose the second, because the mapping table becomes a permanent dependency of every replay path. - Cut consumers over one at a time with a parallel run, comparing projection outputs row by row. The search index and the warehouse go last and take longest, because they hold their own derived copies keyed on fields you are now encrypting.
- Delete the first key only when every consumer is cut over and verified.
Where data can diverge
Between steps 3 and 4 two streams exist and a consumer can be subscribed to the wrong one. The symptom is not an error: it is a projection that is silently a few hours behind on one entity type. Detect it with a continuous count-and-checksum comparison per entity type, not with a sampled spot check.
The point of no return
Destroying the first data key. Everything before it is reversible. After it, the affected events are ciphertext with no key, any consumer that was not cut over produces gaps rather than failures, and there is no restore that brings the plaintext back, which is the whole point. Gate the first destruction behind an explicit sign-off and run it on a single test subject first.
What cannot be migrated
Write-once archives. An object held under S3 Object Lock in compliance mode cannot be overwritten or deleted by any user including the account root, and the retention period cannot be shortened, so those copies survive until they expire on their own schedule. The honest answer to a regulator is therefore a deletion horizon: production within seconds, warehouse within a day, backups within the backup retention, WORM archive at its stated expiry. Put the number in writing rather than claiming immediacy you cannot deliver.
For scale: 8 billion events re-encrypted at a sustained 200,000 events per second is about 11 hours of pure compute. The realistic elapsed time is weeks, because you throttle to avoid disturbing production and because the consumer cutovers, not the cryptography, set the pace.
When not to crypto-shred at all
If the log's retention is already short, the cheaper answer is to stop writing personal data into it and wait the retention out. A 90-day log needs no key registry, no backfill and no key-destruction runbook. Crypto- shredding earns its permanent key-management burden only when retention is measured in years and the log is genuinely the system of record. If the personal fields could instead live in a mutable side store referenced by id from the log, do that: deleting a row is cheaper to build and cheaper to prove than managing 40 million keys.