Consent & Privacy Service

Architecture Views

22 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

Twenty-two views in seven acts, for the platform behind the cookie banner, the preference centre, "download your data" and "delete my account". Read them in order: the set argues one thing, which is that this platform is the system of record for permission and never for personal data, and every later view either honours that boundary or is a consequence of it. Every number in the set is a stated assumption, chosen so a reviewer can disagree with one and follow it to the decision that depends on it.

Context and scope

Who acts on this platform, what it is accountable to, and the deliberate omission at the centre of it: the data itself.

People and journeys

Who it is for, what each of them gets to do, and the three journeys where it either works or is a screen that lies.
03 People whose data it is Everyday user 180 M registered Goal — When I say stop, I want it to actually stop — everywhere, not just on the screen where I said it. Core journeys Withdraw a consent Delete my account Download my data Visitor, not signed in pre-identity Goal — I should be able to refuse before anyone knows who I am, and have that refusal survive signing in. Core journeys Refuse before signing in People accountable for it Privacy officer 3 jurisdictions Goal — Prove what we were allowed to do on a given day without asking six teams and hoping. Core journeys Answer a regulator Purpose owner 40 product teams Goal — Ship a new use of data without becoming the reason we are fined. Core journeys Register a purpose See my purpose's cost Support agent case-bound, time-boxed Goal — Act on what the caller is asking for without being shown their whole life. Core journeys Raise a case for a caller Machines and third parties Product service 40 integrated Goal — Tell me yes or no in under ten milliseconds, and never make me guess when you cannot. Core journeys Ask for a decision Processor 60 recipients Goal — Send me one instruction I can act on, and tell me exactly what you expect back. Core journeys Act on an erasure instruction Supervisory authority statutory deadlines Goal — Show me the record as it stood, not a reconstruction assembled after my letter arrived. Core journeys Request the evidence pack Actors and Their Core Journeys Person or role Journey / task External / third party Three of these actors want the platform to say no. That is the product. v 1.0 · owner Security & Identity Architecture Actors and Their Core Journeys Eight actors, and the uncomfortable observation that three of them are paying for the platform's ability to refuse. HTML page SVG draw.io

Structure

The parts, the regional boundary they are repeated inside, and the contracts that cross it.

Data

What is stored, what is derived, what is irreplaceable, and what is deliberately not here.

Runtime

What actually happens on a withdrawal, on an erasure, and on a read that has to ask permission first.

Operations

How it is deployed per jurisdiction, what is watched, and how a purpose gets into production.

Assurance

Why the platform is not itself the easiest way to attack privacy, and what happens when each of its bounds fails.

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

This platform is the system of record for permission and never for personal data: it holds what each person allowed, and the data stays with its owner, which must ask.

Everyone has used this system without being told its name. The cookie banner. The "manage your ad preferences" screen three taps into an app. The email footer that says unsubscribe and sometimes means it. The "download your data" button that mails a zip the next day. The "delete my account" flow that warns you it is permanent and then has to make it permanent across forty internal systems and sixty vendors that were each sent a copy of you years ago. The screens are easy. What is hard is that each of them is a claim about the behaviour of an entire estate, made to a person who has no way to check it, enforced by a company whose incentive is to use the data, and audited by a regulator who will ask what was permitted on a specific Tuesday eighteen months ago. A privacy programme that cannot answer that question in minutes does not have a privacy capability; it has a set of promises and a document describing them.

Hold permission, not people. Purposes are versioned entities with a lawful basis declared per jurisdiction, a retention period, and the explicit list of systems and recipients permitted to process under them — and that list is the authoritative fan-out for every withdrawal and every erasure, which makes the registry a correctness dependency rather than a catalogue. Every permission decision is appended to a per-region ledger that is never updated; current state is a disposable projection carrying the ledger version it was built from, and every answer the platform gives names that version. Purpose limitation is enforced where the data is read, through a decision API backed by in-process caches with a published staleness ceiling and a per-purpose fail-closed posture. Rights requests become durable, resumable cases that resolve the subject to every identifier they are known by and fan out to roughly a hundred targets, tracking four distinct states per target — instructed, acknowledged, attested, verified — with a suppression list guarding every path by which data gets back in, and sampled absence probing because an attestation is a claim and not a fact. The whole regional stack is repeated per jurisdiction with no cross-border replication and no failover, and only the registry, which holds no personal data, is global.

What it is, and what it is not

The system of record for permissionA central lake that copies personal data in so it can be deleted in one place
Purpose limitation enforced at the point of useFiltering at collection, which cannot honour a later grant
A ledger of decisions, append-onlyA consent table that is updated when someone changes their mind
Erasure as a verified protocol with four states per targetA DELETE with a return code
Silence from a target treated as failureSilence treated as pending, so the case can close
Completeness evidenced and its residual gap measuredCompleteness guaranteed
Fail closed, with a named, approved, auditable set of exceptionsFail open because the decision path had a bad minute
Residency enforced in the data pathResidency asserted in a policy document and a region column
UNKNOWN as a verdict the caller must handleUNKNOWN quietly coerced to ALLOW

The decisions that are the architecture

01Permission is the system of record; personal data stays with its owner

The platform knows what each person allowed. It does not hold their data. That one boundary is why erasure is an orchestrated protocol rather than a delete, why purpose limitation lives on the read path, and why the purpose-to-system map is a correctness dependency.

ADR-01

02The identifier graph is centralised, and is the most constrained store in the design

Proving an erasure reached every identifier requires knowing every identifier, which means building the most attractive target in the company. It is accepted deliberately, held per region, encrypted per subject, case-bound and audited per record read.

ADR-02

03The consent ledger is the truth; current state is a projection

Append-only, never updated, RPO 0. A grant after a withdrawal is two facts. Every current-state record and every decision names the ledger version it came from, which is what makes point-in-time replay a query rather than a reconstruction.

ADR-03

04Lawful basis is per purpose version per jurisdiction, and widening re-asks

The same purpose has different bases in different places, and a change that adds a data category or a recipient is a new purpose version requiring fresh permission. Modelling this as data rather than as review is what stops scope creep being invisible.

ADR-04

05There is no default purpose; an unregistered one is rejected, not denied

DENY for an unknown purpose would let a typo look like a lawful refusal. Rejection surfaces the real problem: something is processing data under a name nobody declared.

ADR-05

06Purpose limitation is enforced at the point of use

Collection-time filtering is cheap, irreversible and cannot honour a later grant. Read-time decision is correct and couples the hottest read paths in the company to this platform, which the next two decisions exist to make survivable.

ADR-06

07Consent state is pushed to local caches with a published staleness ceiling

The common decision is an in-process lookup; the ledger stays authoritative. Staleness becomes the central risk, so it is measured per enforcement point and bounded by a fifteen-minute ceiling whose breach is an incident.

ADR-07

08Fail closed, with fail-open as a declared per-purpose posture

Universal fail-closed makes this platform the ceiling on every product's availability. A declared exception set puts a legal judgement in a configuration field — so it lives in the registry behind two-person approval and is audited whenever it is relied on.

ADR-08

09UNKNOWN is a distinct verdict the caller must handle

Collapsing "no record" and "could not read the record" into one answer is how a broken read path becomes silent unlawful processing. The SDK contract test, not a promise, is what enforces it.

ADR-09

10Erasure is a four-state verified protocol per target

Instructed, acknowledged, attested, verified. Collapsing them into a boolean is how a case closes while data remains, and it is the single most common way a privacy programme is wrong about itself.

ADR-10

11The erasure technique is declared per store, in advance

Hard delete, crypto-shredding a per-subject key, irreversible anonymisation, or tombstone-plus-suppression. Immutable logs and columnar warehouses rule out some options, so the technique is registry data and part of the evidence.

ADR-11

12Silence from a target is failure, not pending

A target that neither acknowledges nor attests inside its window blocks the case and escalates to its named owner. The alternative — pending until a timer closes it — manufactures completed erasures that never happened.

ADR-12

13The suppression list is mandatory on every ingestion and restore path

Backups, replayed streams and vendor re-uploads will reintroduce erased subjects. One compact set of hashed keys, consulted everywhere data enters, is the only control that holds across all of them.

ADR-13

14Residency is enforced in the data path, and consent metadata is personal data

Regional stacks, per-region keys, no cross-border replica and no failover. A region's unavailability denies its own subjects' consent-based purposes, which is correct rather than a gap.

ADR-14

15Evidence of an erasure survives the erasure

The platform retains the minimum identifiers needed to prove a deletion happened, under a legal-obligation basis the subject cannot withdraw, and says so to the subject plainly instead of engineering the tension away.

ADR-15

Why this should still be right in ten years

The regulations will change, the cloud services will be renamed, and the consent UX will be redesigned twice. The parts of this design that should outlive all of that are the ones that are statements about where authority sits and what can be proved, rather than statements about products or statutes.

Permission and data are different kinds of thing

A system that governs use without holding the data has a different failure mode from one that centralises the data to govern it. That distinction is older than GDPR and will outlast whatever replaces it; the specific rights, deadlines and bases are configuration on top of it.

An append-only record of what was permitted

Every regulator, auditor, court and angry user asks the same question in the same shape: what were you allowed to do, and when. A ledger answers it; a mutable state table never will, however carefully it is backed up.

Completeness is evidenced, not guaranteed

No platform can prove by inspection that another system deleted something. Designs that claim otherwise are claiming a guarantee they cannot hold. Measuring and publishing the residual gap is the only honest posture, and it does not depend on which systems are in the estate.

The fan-out list is a correctness dependency

The hard part of deletion has never been deleting. It is knowing where the copies are. Any future version of this problem is still a registry-completeness problem wearing new technology.

Residency is a property of the data path

Jurisdictional boundaries are getting stricter and more numerous, not fewer. A design whose boundary is enforced by routing and keys survives new jurisdictions as new deployments; a design whose boundary is a policy document does not survive the first audit.

Non-functional targets

Every target below is a stated assumption for this exercise, chosen so a reviewer can disagree with one and follow it to the decision that depends on it.

QualityTargetHow it is metView
Decision path availability ≥ 99.99% monthly per region In-process cache first; regional stack across three AZs with no global request-path dependency 16
Consent capture availability ≥ 99.95% monthly Durable commit to a regional quorum before acknowledgement; never degrades to no-record 02
Decision latency p50 ≤ 2 ms, p99 ≤ 10 ms on cache hit; p99 ≤ 40 ms remote Version-stamped local cache refreshed from the withdrawal stream and the projection 15
Capture latency p99 ≤ 200 ms, durably committed Single append to the ledger; no wait on projection or propagation 13
Withdrawal propagation p95 ≤ 5 s, p99 ≤ 60 s, ceiling 15 min Withdrawal event stream per subject, lag measured per enforcement point 10
Recipient propagation ≤ 24 h programmatic, ≤ 7 days manual default Strongest mechanism each recipient supports, recorded on the case 14
Rights acknowledgement ≤ 24 h, with a case identifier Durable intake independent of the fan-out 05
Erasure completion p95 ≤ 7 days first-party, ≤ 30 days absolute Four-state per-target tracking with escalation at the declared window 14
Decision throughput 120,000/s steady, 400,000/s for 120 s Read path scaled independently of capture; batch path on immutable snapshots 02
Capture throughput 4,000/s steady, 200,000/s in a withdrawal storm Durable queueing of propagation with per-subject ordering; capture stays in budget 10
Durability RPO 0 ledger, audit and keys; RPO 60 s projections Regional quorum commit; projections rebuilt from the ledger 11
Recovery RTO 10 min decision path, 60 min full projection Replay at ≥ 20× real time; no cross-jurisdiction failover offered 11
Late-ALLOW correctness Zero tolerated beyond the 15 min ceiling Each occurrence a reportable defect with a named owner, not a percentage 22
Erasure detection confidence ≥ 99% of targets with ≥ 1% residual within 7 days Budgeted sampled absence probing with a declared sample rate 14
Evidence reproducibility 100% of subjects inside retention Point-in-time ledger replay including the notice version shown at capture 06

Scope

In scope

  • Purpose and lawful-basis registry with versioning, per-jurisdiction basis, data categories, retention periods, and the purpose-to-system and purpose-to-recipient maps
  • Consent capture for authenticated and unauthenticated subjects, withdrawal at parity with grant, and the strictest-wins merge on sign-in
  • Append-only regional consent ledger with a rebuildable, version-stamped current-state projection
  • Decision API with in-process last-known-good caching, UNKNOWN as a distinct verdict, and a per-purpose failure posture
  • Withdrawal event stream with propagation lag measured per enforcement point against a published ceiling
  • Rights cases for access, portability, rectification, restriction, objection and erasure, with proportionate verification and durable resumable orchestration
  • Erasure fan-out with four states per target, declared per-store techniques, a mandatory suppression list, and sampled absence verification
  • Residency enforced in the data path with per-region keys and no cross-border replication or failover
  • Append-only tamper-evident audit, point-in-time consent reproduction, and records-of-processing generation from live configuration

Explicitly out of scope

  • Being the store of personal data itself — the estate keeps its own data and asks
  • Authentication and account management; the platform consumes identity and does not issue it
  • Campaign execution, advertising delivery and the consent banner's front-end implementation
  • Deciding which lawful basis a purpose has; counsel decides, the platform enforces and evidences
  • Data discovery and classification across the estate — the purpose-to-system map is declared, and its incompleteness is a named risk
  • Model retraining itself; the platform triggers it and records the exemption where a subject is not identifiable

What a four-week prototype should prove

The prototype's job is to falsify the three claims the design rests on: that a withdrawal reaches a realistic population of enforcement points inside the published ceiling while the system is busy, that a decision on the read path is cheap enough that product teams will actually call it, and that an erasure across heterogeneous stores can be verified rather than merely attested. Everything else in this package is a consequence of those three holding.

  1. One region, one purpose registry holding 12 purposes across 3 jurisdictions, and 2,000,000 synthetic subjects.
  2. Six enforcement points with real in-process caches, two of them deliberately pinned to an older SDK version.
  3. Four target classes behind real adapters: a transactional table, an append-only log with per-subject keys, a columnar warehouse, and a mock third-party API that goes silent at random.
  4. A ledger preloaded with 90 days of synthetic capture history, so replay is measured against a realistic partition rather than an empty one.
  • Measure withdrawal propagation per enforcement point under a 50x capture burst and report the distribution rather than the mean: the ceiling claim fails if the slowest point is outside 15 minutes
  • Put the decision call in the read path of one genuine product surface and measure the added p99; if it is not single-digit milliseconds on cache hit, read-time enforcement will be negotiated away in production
  • Erase 10,000 subjects across all four target classes and verify absence independently, counting what the attestations claimed against what the probes found
  • Rebuild the full current-state projection from the ledger and measure the replay rate, because the 60-minute RTO is entirely a bet on 20x real time or better
  • Switch the regional decision API off and confirm the fail-closed posture actually holds at every enforcement point, including the two running the older SDK
  • Run one point-in-time evidence query for a subject at an arbitrary past timestamp, including the notice text shown at capture, and time the whole path end to end

Open risks, carried rather than hidden

RiskIf it landsResponse
The purpose-to-system map is incomplete, because it is declared by the teams it constrains Withdrawal and erasure fan out to a list that is missing systems, producing erasures that look complete and are not — the design's worst failure, and a silent one Unregistered-use findings from declared purpose on every decision call, periodic estate-wide re-identification scanning in Phase 3, and a named accountable owner per purpose whose findings are tracked rather than closed
The fifteen-minute propagation ceiling is a claim nobody tests until the first real withdrawal storm Unlawful processing accrues across forty enforcement points at once, and the first evidence of it is a regulator's letter Propagation lag reported per enforcement point as a routine compliance signal, a paging alarm on ceiling breach, and deliberate burst drills in production rather than only in load tests
Read-time enforcement is negotiated away under latency pressure Purpose limitation silently reverts to collection-time filtering, which cannot honour a later grant, and the whole decision plane becomes decorative Single-digit-millisecond cache-hit budget treated as a product requirement, the remote-evaluation ratio monitored, and a client SDK that makes the correct call the easy one
The centralised identifier graph is compromised An attacker learns the linkage between every pseudonym and every person in the estate — worse than any single data breach the platform was built to limit Separate authorisation from the decision path, per-subject encryption, case-bound just-in-time access, per-record read auditing, and acceptance that this store is the design's largest residual risk
A single signed policy bundle reaches nine regions in two minutes One bad approval is a global misconfiguration of lawful basis, retention or fail-open posture Two-person approval, counsel sign-off, staged regional adoption, and a tested rollback to the previous bundle version measured in minutes
Manual recipients and legally retained categories accumulate as permanent open gaps The compliance dashboard normalises a tail of never-closing obligations, and the organisation stops reading it Refuse at registration any purpose whose recipient cannot meet the jurisdiction's window, assign every manual obligation an owner and a due date, and report the tail's size as a tracked number rather than a list

Architecture Decision Record

Why every component and every technology on these 22 views is what it is, and what each choice costs.

Fifteen decisions make up this architecture. Everything else across the twenty-two views is a consequence of one of them, and each record carries the question that forced it, how it is realised on AWS, the alternatives including the ones that are right for a different organisation, and the conditions under which the choice should be reversed.

Status of this document. This is a design, not a report on a running system. Every rate, latency, volume, retention and cost figure in this package is a stated assumption, chosen to be defensible for a consumer SaaS of roughly 180 million registered subjects across 40 internal processing systems, 60 external recipients, 9 regional deployments and 3 jurisdictional boundaries, carrying 120,000 permission decisions a second at steady state and 25,000 rights requests a day. They are written as numbers so that they can be argued with and corrected, which vagueness does not allow.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

The boundary 2

What this platform owns, and what it deliberately refuses to own.

ADR-01Permission is the system of record; personal data stays with its owner ADR-02The identifier graph is centralised, and treated as the most dangerous store in the design

Permission as data 3

How a permission is recorded, versioned, and made provable eighteen months later.

ADR-03The consent ledger is the system of record; current state is a disposable projection ADR-04Lawful basis is declared per purpose version per jurisdiction, and widening re-asks ADR-05There is no default purpose, and an unregistered purpose is rejected rather than denied

Enforcement 4

Where the purpose question is asked, and what the answer is when it cannot be answered.

ADR-06Purpose limitation is enforced at the point of use, not at collection ADR-07Consent state is pushed to local caches with a published staleness ceiling ADR-08Fail closed by default, with fail-open as a declared per-purpose posture ADR-09UNKNOWN is a distinct verdict, and a caller may never read it as ALLOW

Erasure as a protocol 4

Why deletion is a distributed protocol with verification rather than a command with a return code.

ADR-10Erasure is a four-state verified protocol per target, never a boolean ADR-11The erasure technique is declared per store, in advance, and is part of the evidence ADR-12Silence from a target is failure, not pending ADR-13The suppression list is mandatory on every ingestion and restore path

Residency and evidence 2

Where the record of permission lives, and what has to survive the request to be forgotten.

ADR-14Residency is enforced in the data path, and consent metadata is personal data ADR-15Evidence that an erasure happened survives the erasure

Technology by capability

Every capability below is satisfied by an AWS service or an open-source component chosen for a property the requirement names, not for familiarity. Where a simpler or more obvious option was rejected, the reason is the architecture rather than the service.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Consent ledger (system of record) DynamoDB, keyed on subject_key with captured_at sort key, conditional writes forbidding overwrite, Streams enabled Amazon Web Services Aurora PostgreSQL append-only table; Kinesis as the log A subject key is naturally high-cardinality, so partitioning by it removes hot keys entirely at 2.5 billion current-state records, and Streams gives the projection builder a durable ordered feed without a second system. ADR-03
Current-state projection DynamoDB single-item-per-(subject, purpose, jurisdiction), version-stamped, rebuilt from the ledger Amazon Web Services ElastiCache as the primary store; a relational materialised view The read is a point lookup at 120,000/second with no range scan, and the store must be rebuildable rather than durable — which rules out anything whose contents cannot be discarded. ADR-03
Purpose and lawful-basis registry Aurora Global Database, strongly consistent in the primary region, globally readable Amazon Web Services DynamoDB global tables; a Git repository as the sole source The registry is small, highly relational (purpose, version, basis, target), and needs real constraints and transactions; it holds no personal data, which is the only reason it is allowed to be global at all. ADR-04
Policy distribution to regions and SDKs Signed bundle in S3 fronted by CloudFront, versioned, verified on load Amazon Web Services Direct reads from the registry; AppConfig A decision plane must keep working with the global registry unreachable, so it needs a local artefact with a version it can report — and a signature so a cached bundle cannot be substituted. ADR-07
Withdrawal fan-out EventBridge, one bus per region, per-subject ordering, one subscription per enforcement point Amazon Web Services SNS fan-out; Kafka; polling the projection Per-enforcement-point subscriptions are what make per-point propagation lag measurable, which is the metric the whole 15-minute ceiling claim depends on. ADR-07
Rights-case orchestration Step Functions standard workflows, one execution per case, DynamoDB for target outcomes Amazon Web Services A homegrown saga on SQS; Temporal A case runs for up to thirty days across a hundred targets with retries, timers, escalations and human refusal steps; durable resumable execution with per-step history is the requirement, and the history doubles as evidence. ADR-10
Identity resolution index DynamoDB table with per-item encryption under per-subject KMS data keys, separate authorisation, per-record read audit Amazon Web Services A graph database; resolution assembled per case from source systems A graph engine's traversal power is not needed for alias lookup and would widen what a compromise yields; per-subject keys make the linkage disappear when the subject is crypto-shredded. ADR-02
Crypto-shredding KMS data key per subject per region, wrapped by a regional CMK with no cross-region grant Amazon Web Services Application-managed keys in a secrets store; envelope keys per tenant Erasure from append-only stores has to be key destruction, and the key boundary has to coincide exactly with the residency boundary, which a regional CMK with no grant enforces structurally. ADR-11
Tamper-evident evidence store S3 with Object Lock in compliance mode, hash-chained records, independent credentials Amazon Web Services QLDB; an append-only table in the operational database The evidence must survive compromise of the operational plane, so it needs a separate credential domain and a retention mode that the platform's own administrators cannot shorten. ADR-15
Batch and campaign-scale evaluation Immutable projection snapshot exported to S3, queried with Athena Amazon Web Services Looping the per-request decision API; a read replica A million-subject eligibility question must not share capacity with the read path, and a snapshot gives a single consistent version to report alongside the result. ADR-06
Compute for capture, decision and adapters ECS on Fargate across three AZs per region, with the decision SDK in-process in consuming services Amazon Web Services Lambda; EKS Steady high-throughput services with warm caches are a poor fit for per-invocation isolation, and EKS adds a control plane this platform does not need for a dozen services per region. ADR-01
Subject authentication and step-up Cognito for subject identity with step-up for export and erasure; workforce IdP federation for agents Amazon Web Services The product's own session as sufficient proof An unverified erasure request is an attack on the subject, so the verification strength has to be a property of the request type rather than of the session that happens to exist. ADR-02
Suppression list distribution DynamoDB per region plus a compact snapshot on S3 and CloudFront for local copies Amazon Web Services A shared database table; a published file per consumer Every ingestion and restore path must be able to consult it, including batch paths with no network access to the platform's APIs — so it needs both an online and an offline form. ADR-13
Observability CloudWatch metrics dimensioned per enforcement point and per purpose, X-Ray on the decision path, findings routed to owning teams Amazon Web Services Platform-level aggregate dashboards Propagation lag and UNKNOWN rate are only meaningful per enforcement point; a platform average hides the single cache that stopped consuming, which is the failure that matters. ADR-12

The decisions, and the alternatives that lost

The boundaryWhat this platform owns, and what it deliberately refuses to own.

ADR-01

Permission is the system of record; personal data stays with its owner

Accepted

To govern how personal data is used, should the governing platform hold that data, or only the record of what may be done with it?

Context
There is an obvious design that makes everything else easy: copy the personal data into one governed store, and then consent, retention, export and erasure are all local operations on a single schema. Erasure becomes a delete you can prove by inspection. Export becomes a query. Purpose limitation becomes a column filter. The cost is that a platform whose purpose is to reduce the harm of holding data about people becomes the largest aggregation of data about people in the company, with its own residency, retention, consent and breach story for the copy, and a freshness problem for every field that moved. The alternative keeps the data where it is and makes the platform authoritative only over permission — which pushes the difficulty outward into forty systems and sixty vendors that must each be instructed, and must each be believed or verified.
Decision
The platform is the system of record for permission and never for personal data. It holds purposes, lawful bases, consent decisions, identifier linkage, case state and evidence. Personal data remains in the systems that own it, and those systems must ask before use. Every capability the platform offers is therefore either a decision about permission or an orchestration across systems it does not control.
How it is realised on AWS
The regional plane holds the consent ledger and current-state projection in DynamoDB, the identity index in a separately authorised DynamoDB table with per-subject KMS keys, and rights-case state in Step Functions plus DynamoDB. No product data is replicated in. Consuming systems integrate three ways: an in-process SDK that calls the decision API, an EventBridge subscription for withdrawals, and an adapter for erasure instructions. The purpose registry in Aurora Global holds the authoritative list of which systems and recipients each purpose touches.
Options weighed
  • ChosenPermission as system of record; data federated with its owner: Keeps the blast radius of a compromise to permission metadata; pays with orchestration, verification and an incomplete-map risk.
  • RejectedCentral governed copy of personal data: Provable single-point erasure, and it builds the thing the platform exists to limit, with freshness and residency problems for every copied field.
  • RejectedPure policy-decision point with no state at all: Cannot answer what was permitted last March, cannot run a rights case, and cannot prove anything to a regulator.
  • Right elsewhereEmbedded governance in each system, no central plane: Right for a small estate with one data store; at forty systems it guarantees forty inconsistent interpretations of the same consent.
Consequences
What it buys
  • A compromise of this platform yields what people refused, not who they are and what they did — a materially smaller harm than the central-copy alternative.
  • Each owning system keeps responsibility for its own data, which is where the knowledge of how to delete it correctly actually lives.
  • The platform's own residency, retention and consent story stays small enough to reason about.
What it costs
  • Erasure can never be proved by inspection. It is instructed, attested and sampled, and the residual gap is a measured quantity rather than zero.
  • The purpose-to-system map becomes a correctness dependency: a system missing from it is a system that will not be told.
  • Purpose limitation must be enforced on other teams' read paths, which makes adoption a product problem as much as an architectural one.
Choose differently when
If the estate consolidated onto one or two data platforms with a shared access layer, the central-copy design's costs mostly disappear and its provability becomes worth having. The decision should also be revisited if regulators begin to require demonstrated rather than evidenced completeness, because no federated design can supply that.
Why it holds up over time
This is a statement about what kind of thing the platform is authoritative over. It survives replacing every service named here, every regulation cited, and the entire rights catalogue, because none of those changes whether permission and data are the same object.
LessonDecide what your platform is the truth about before deciding what it stores. A governance system that centralises the thing it governs usually inverts its own goal.
Shown on views01 02 08
ADR-02

The identifier graph is centralised, and treated as the most dangerous store in the design

Accepted

To reach every copy of a person, something must know every identifier they are known by. Should that knowledge be centralised, federated, or assembled per request and discarded?

Context
A subject is a user id in one system, a device identifier in another, a hashed email at a vendor, a customer number in billing, and three anonymous session identifiers that were later linked. An erasure that reaches one of those is a failure that looks exactly like a success. Completeness therefore requires a resolution step, and the question is where the linkage lives. Central resolution is the only option that can demonstrate completeness, and it builds a store that maps every pseudonym to every person — the single most valuable target in the company, created by the team whose job is reducing privacy risk. Federated resolution avoids that store and makes completeness unprovable: each system answers for itself, and nobody can say whether the set was whole.
Decision
The identifier graph is centralised, per region, as a first-class store — and is then treated as the most constrained component in the architecture. It is authorised separately from the decision path, encrypted with per-subject keys, queryable only in the context of an open rights case or a decision evaluation, and audited at individual-record granularity. Holding it is accepted as the design's largest residual risk, stated rather than mitigated away.
How it is realised on AWS
A dedicated DynamoDB table holds identifier hashes mapped to subject keys, with items encrypted under the subject's own KMS data key so that destroying that key removes the linkage as well as the data. Access requires a case-bound role issued for a specific case id with a 30-minute expiry; the resolver service is the only principal with read access and emits one audit record per item returned to S3 Object Lock. Decision-path credentials cannot read the table at all.
Options weighed
  • ChosenCentral graph, maximally constrained and audited: The only design that can evidence completeness; concentrates risk in one store and says so.
  • RejectedFederated resolution, each system answers for itself: No honeypot, and no way to state whether an erasure reached everything — which makes every completion claim unfalsifiable.
  • RejectedAssembled per case from live systems, then discarded: Attractive until a case must be reopened or audited, at which point the evidence of what was resolved no longer exists.
  • Right elsewhereNo resolution; subject must name their own identifiers: Defensible for a single-product company with one identifier; at this scale it pushes an impossible task onto the person least equipped for it.
Consequences
What it buys
  • An erasure case can state which identifiers it covered, which makes a reopened case and an audit both tractable.
  • Per-subject encryption means crypto-shredding removes the linkage, so the graph does not outlive the subjects in it.
  • Per-record read auditing is the only evidence that the graph was not browsed, and it exists from day one rather than after an incident.
What it costs
  • The company now holds a store whose compromise is worse than most of the breaches the platform was built to limit.
  • Per-record audit at case volume is a real cost that will be questioned in the first budget review.
  • Linking identifiers is itself processing, and needs its own lawful basis and its own entry in the records of processing.
Choose differently when
If the estate adopted a single pseudonymous subject identifier end to end, the graph shrinks to a trivial mapping and most of this record's cost disappears. Conversely, if a jurisdiction holds that maintaining such a graph is unlawful irrespective of purpose, resolution must become federated and the completeness claim must be weakened in public.
Why it holds up over time
The need to resolve a person to their aliases is intrinsic to acting on their behalf across systems, and no technology removes it. What may change is whether that resolution is allowed to be stored, which is the one condition that would force a redesign.
LessonWhen a privacy control requires building a privacy risk, say so in the architecture rather than in a footnote. A risk that is named gets constrained; a risk that is implied gets reused.
Shown on views12 20 21

Permission as dataHow a permission is recorded, versioned, and made provable eighteen months later.

ADR-03

The consent ledger is the system of record; current state is a disposable projection

Accepted

Should a person's current permission be stored as a row that is updated when they change their mind, or derived from an append-only record of every decision they ever made?

Context
A consent table is the natural implementation and it loses the question everybody eventually asks. "Were we allowed to send that email on 14 March" cannot be answered from a row that was overwritten in April, and a backup restore answers it with whatever happened to be in last night's snapshot. An append-only ledger answers it exactly, at the cost of a derived read path, a rebuild story, and an eventual-consistency window between a capture and the state every reader sees.
Decision
Every permission decision is appended to a per-region, per-subject-partitioned ledger and is never updated or deleted inside its retention period. Corrections are compensating entries. The current-state projection is derived, rebuildable, and carries the ledger version it was built from; every decision the platform returns names that version. RPO 0 for the ledger, RPO 60 s for the projection.
How it is realised on AWS
The ledger is a DynamoDB table keyed on subject_key with a sort key of captured_at, written with a condition that forbids overwrite, and acknowledged only after a regional quorum commit. DynamoDB Streams feeds a Fargate projection builder that maintains the current-state table and emits withdrawal events to EventBridge. A full projection rebuild runs from the stream's retained history and the ledger partition scan; archive beyond the hot window lands in S3 and is replayable.
Options weighed
  • ChosenAppend-only ledger as SoR, state as projection: Point-in-time truth, free audit trail, rebuildable reads; pays with a derived read path and a staleness window.
  • RejectedMutable consent table with an audit trigger: Simpler and immediately consistent, and the audit trail is a side-effect nobody tests until it is the only evidence.
  • RejectedEvent sourcing with no materialised state at all: Correct and far too slow for a 120,000/second read path; the projection is not optional at this scale.
  • Right elsewhereBitemporal relational model: Right where the whole estate is already relational and volumes are modest; at billions of current-state records the write amplification is unattractive.
Consequences
What it buys
  • Point-in-time reproduction is a query rather than a reconstruction, which is what removes the trough from the regulator journey.
  • A bad projection deploy, a corrupt partition or a wrong derivation rule is repaired by rebuilding, not by surgery on live consent.
  • A grant after a withdrawal is two facts, so a dispute about sequence is answerable.
What it costs
  • Readers see state that may be up to 60 seconds behind a capture, which is why the decision response carries a version.
  • The 60-minute projection RTO rests entirely on achieving ≥ 20× real-time replay, which is an assumption and not a measurement.
  • Ledger growth is unbounded by design — 40 TB over seven years on the stated assumptions — and needs tiering rather than pruning.
Choose differently when
If retention requirements collapsed to months rather than years and nobody needed point-in-time answers, a mutable table with change data capture would be cheaper and sufficient. The decision should also be revisited if replay cannot be made fast enough to hold a credible RTO, because then the projection has quietly become the system of record whatever the diagram says.
Why it holds up over time
"The record of decisions is the truth and the current state is a view of it" is a modelling choice that survives every storage technology. It is the same decision an accounting ledger made, for the same reason.
LessonIf anyone will ever ask what the state was, store the decisions and derive the state. Retrofitting history onto a mutable table is not possible, only approximated.
Shown on views11 12 13
ADR-04

Lawful basis is declared per purpose version per jurisdiction, and widening re-asks

Accepted

Is a purpose one thing with one legal justification, or a versioned thing whose justification depends on where the subject lives?

Context
The same processing — recommending content from behaviour — is consent in one jurisdiction, legitimate interest in another, and arguably contractual necessity in a third. Modelling basis as a property of the purpose forces one of three bad outcomes: the strictest basis applied everywhere, which forfeits lawful processing; the loosest applied everywhere, which is unlawful; or a purpose duplicated per jurisdiction, which makes a withdrawal in one place silently fail to cover the same processing described under another name. Separately, purposes change. Adding a data category or a recipient to an existing purpose is a different thing from narrowing it, and treating both as an edit means a grant given for something small silently covers something larger.
Decision
Lawful basis is an attribute of the triple (purpose, version, jurisdiction) and exactly one basis is declared per triple. Every purpose change produces a new version. A change that widens scope — a new data category, a new recipient, a longer retention period — is flagged as widening and requires fresh permission from subjects whose basis is consent; a narrowing change applies immediately to existing permissions and does not re-ask.
How it is realised on AWS
purpose, purpose_version and lawful_basis are three tables in Aurora Global, with basis keyed on (purpose_id, version, jurisdiction) and carrying the notice version to show. The widening flag is set in reviewed policy-as-code and cannot be set by the proposing team alone; the release path gates on it. Regional decision planes evaluate basis from the signed bundle, so a jurisdiction's rule is applied by the plane that serves that jurisdiction.
Options weighed
  • ChosenBasis per (purpose, version, jurisdiction); widening re-asks: Expresses reality; makes scope creep visible as a flag with an approval, rather than invisible as an edit.
  • RejectedOne basis per purpose, strictest wins globally: Safe, simple, and forfeits lawful processing the organisation is entitled to — which guarantees the platform is bypassed.
  • RejectedA separate purpose per jurisdiction: Looks tidy until a withdrawal covers one purpose and not its twin, which is the failure mode hardest to detect.
  • Right elsewhereBasis decided at evaluation time from a rules engine: Right where jurisdictions change faster than releases; here it makes the basis unprovable after the fact, which defeats the evidence requirement.
Consequences
What it buys
  • A decision response can name the basis it relied on, which is the first thing a regulator asks after the verdict.
  • Scope creep becomes a visible flag with two approvers rather than an unremarkable configuration edit.
  • A new jurisdiction is a new basis row and a new regional deployment, not a schema change.
What it costs
  • Re-asking is unpopular, and the pressure to classify a widening change as narrowing will be constant and will come from people with good intentions.
  • An expiry or re-ask wave concentrates prompts and looks to subjects like the banner returning, which needs operational spreading the model does not provide.
  • Purpose version proliferation makes the registry harder to read than a single row per purpose would be.
Choose differently when
If the organisation operated in a single jurisdiction, basis collapses to a purpose attribute and the versioning can be simplified. If regulators converged on a common basis taxonomy and a common test for what counts as widening, the flag could become derived rather than declared.
Why it holds up over time
Jurisdictional divergence in data protection is increasing, not decreasing. A model that treats basis as varying by place and time will fit the next decade better than one that treats it as a property of the processing.
LessonModel the dimension your obligations actually vary along, even when today's estate has only one value in it. Adding a dimension later means rewriting every record you already relied on.
Shown on views12 18 19
ADR-05

There is no default purpose, and an unregistered purpose is rejected rather than denied

Accepted

When a caller asks about a purpose the registry has never heard of, is the right answer DENY or an error?

Context
DENY is the tempting answer because it is safe: nothing is permitted, no data is used, and the caller handles it like any other refusal. It is also the answer that hides the only interesting fact in the situation — that something in the estate is processing data under a name nobody declared, which is either a typo that will silently suppress lawful processing or a genuine undeclared purpose, and those need opposite responses. A registry with a permissive default is worse again: it turns the absence of a declaration into permission.
Decision
The registry has no default purpose and refuses to register one that lacks a description, data categories or a retention period. A decision call naming an unregistered purpose returns a typed caller error, not a verdict, and raises an unregistered-use finding against the calling system. Processing cannot proceed, but the caller is told it asked a bad question rather than that it received a lawful refusal.
How it is realised on AWS
The decision API validates the purpose identifier against the signed bundle before evaluating anything and returns a 4xx with a distinct error code. The completeness gate in the release path rejects an incomplete purpose definition at review time. Unregistered-use findings are emitted to CloudWatch with the calling workload identity attached and routed to that system's owner.
Options weighed
  • ChosenTyped rejection plus a finding: Surfaces the real problem and still blocks processing; costs callers one more error path to handle.
  • RejectedReturn DENY: Safe and silent. A typo becomes an unexplained drop in a feature's reach, and an undeclared purpose becomes invisible.
  • RejectedPermissive default for unregistered purposes: Turns an omission into a permission; the opposite of what the platform is for.
  • Right elsewhereAuto-register on first use, pending review: Attractive for a fast-moving startup wanting adoption; here it means production processing under an unapproved basis, which is the thing being prevented.
Consequences
What it buys
  • An undeclared purpose is discovered from traffic rather than from an audit, which is the weakest detection in the package and this is most of it.
  • A typo fails loudly in integration testing instead of quietly in production.
  • The registry stays a complete statement of what the company does, which is what makes records-of-processing generation possible.
What it costs
  • Callers must handle three verdicts and an error, which is a contract the SDK has to make easy or teams will collapse it.
  • A rejected call blocks a feature, so the first production occurrence will be experienced as the platform breaking something.
  • Detection depends on the system declaring the purpose it is processing under; a path nobody instrumented is still invisible.
Choose differently when
If the estate reached a point where purpose declaration was enforced structurally — data access impossible without a purpose token — the finding becomes redundant because the call could not be made. Until then this is the only signal.
Why it holds up over time
The distinction between "not permitted" and "not a question I can answer" is a general API-design truth and will outlast this domain entirely.
LessonNever let a malformed question return a well-formed answer. Collapsing the two hides the only diagnosis worth having.
Shown on views15 19

EnforcementWhere the purpose question is asked, and what the answer is when it cannot be answered.

ADR-06

Purpose limitation is enforced at the point of use, not at collection

Accepted

Where is a purpose actually limited: by never collecting the data, by checking at every read, or by tagging rows and filtering at query time?

Context
Collection-time filtering is the cheapest and most irreversible: if the data was never stored, no purpose can misuse it, and no runtime dependency exists. It also cannot honour a later grant — a subject who opts in next month gets nothing, because the history was discarded. Row tagging in the data plane is the most elegant and requires every query engine in the estate to understand and respect the tags, which is a programme rather than a design. Read-time decision is correct for every case and makes this platform a synchronous dependency of the hottest read paths in the company.
Decision
Purpose limitation is enforced where data is read, through a decision call on the read path, with the result carrying the ledger version it was computed from. Collection remains governed by the registry's data categories, but filtering at collection is not the enforcement mechanism. Tag-and-filter in the data plane is adopted only for the analytics warehouse, where the query engine is one system under one owner.
How it is realised on AWS
A thin client SDK embedded in each of the 40 consuming services evaluates from an in-process cache, falling back to the regional decision API. The call sits inside the service's own read path rather than at its edge, so a purpose question is asked where the data is actually used. The warehouse receives purpose tags plus the suppression feed and filters at query time, which is the one place in the estate where that is cheaper than asking.
Options weighed
  • ChosenRead-time decision at the point of use: Correct in every case including later grants; couples hot read paths to this platform's availability and latency.
  • RejectedCollection-time filtering: Cheapest and safest, and permanently forfeits data a subject later consents to — an irreversible answer to a reversible question.
  • RejectedRow-level purpose tags filtered in the data plane: Best end state and a multi-year programme across every engine; adopted only for the single-owner warehouse.
  • Right elsewherePeriodic batch purge of data whose purpose lapsed: Adequate where reads are infrequent and latency-insensitive; here it leaves a window in which unlawful reads are routine.
Consequences
What it buys
  • A later grant is honoured, because the data still exists and the decision changes.
  • A withdrawal takes effect on the next read rather than on the next purge cycle.
  • The enforcement point is where the context is, so a denial can be handled gracefully by the feature that knows what to do instead.
What it costs
  • Every hot read path in the company now depends on a decision answer, which makes the next two records mandatory rather than optional.
  • Forty integrations means the real enforcement behaviour is as good as the least-updated client library.
  • Teams will ask to cache the verdict for longer than the staleness ceiling allows, and each request will be individually reasonable.
Choose differently when
If the estate consolidated reads behind one or two access layers, tag-and-filter in the data plane becomes achievable and is strictly better: no runtime dependency, no stale verdicts. If the decision call cannot be made cheap enough to survive the hottest path, read-time enforcement will be negotiated away in practice whatever the architecture says, and the honest response is to move to tagging rather than to pretend.
Why it holds up over time
The choice between limiting at collection and limiting at use is independent of technology and is the central design question of any data governance system. The answer here depends on whether consent is reversible, and it is.
LessonAn irreversible enforcement mechanism cannot implement a reversible policy. If the subject can change their mind, the control has to be evaluated, not applied once.
Shown on views10 15
ADR-07

Consent state is pushed to local caches with a published staleness ceiling

Accepted

Should every enforcement point hold a local copy of consent state, or ask the platform for every decision?

Context
Pull gives one authoritative answer and makes this platform a synchronous dependency of 120,000 reads a second — its availability becomes the ceiling on every product's availability, and its worst day is everybody's worst day. Push gives every enforcement point a fast local answer and moves the entire risk into staleness: for some bounded period, a service will cheerfully say yes to something the subject has already refused, which is not a performance defect but unlawful processing. The hybrid that most systems land on — local cache, remote fallback — only works if the staleness is measured where it actually occurs rather than assumed to be small.
Decision
Consent state is pushed. Each enforcement point maintains an in-process last-known-good cache, invalidated by the withdrawal event stream and refreshed from the projection, and every verdict it returns carries the ledger version it was computed from and its own age. Staleness is measured per enforcement point, not as a platform average, and is bounded by a published fifteen-minute ceiling whose breach is an incident with a notification path rather than a line on a dashboard.
How it is realised on AWS
The ledger's stream drives EventBridge, which fans withdrawal events to each consuming service's subscription; the SDK drops the affected subject's cached entries on receipt and re-evaluates on next use. Each SDK instance reports its oldest unrefreshed entry as a CloudWatch metric dimensioned by enforcement point, and a breach of the ceiling pages the platform on-call and the owning team together.
Options weighed
  • ChosenPush with per-point staleness measurement and a published ceiling: Decouples product availability from this platform; makes staleness the one risk and makes it visible where it happens.
  • RejectedPull a decision per read: Always current, and it makes a consent service a hard dependency of every product read path in the company.
  • RejectedShort-TTL cache with no invalidation: Looks like push and is not: the TTL is the propagation guarantee, so the ceiling becomes the TTL and cannot be improved by anything.
  • Right elsewhereCentral enforcement proxy holding state for callers: Right where client libraries cannot be changed; adds a hop and a tier that must be more available than the thing it fronts.
Consequences
What it buys
  • The decision is an in-process lookup at p99 ≤ 10 ms, which keeps it out of every latency budget in the estate.
  • A regional decision-plane outage becomes a bounded staleness event rather than a company-wide failure.
  • Measuring per point exposes the one cache that stopped consuming three hours ago, which an average would hide completely.
What it costs
  • There is a window in which a service says yes to something already refused, and it is a compliance exposure rather than a latency statistic.
  • Every safety guard has to exist in the client as well as the platform, and the clients are owned by other teams.
  • The ceiling, not the event, is the actual guarantee — a lost invalidation is corrected only when the entry ages out.
Choose differently when
If the organisation's read volume were low enough that a central decision service could credibly be more available than every caller, pull is simpler and strictly more correct. The decision should also be revisited if client library heterogeneity becomes unmanageable, because this design silently delegates compliance behaviour to those libraries.
Why it holds up over time
This is a statement about where a copy of the truth is allowed to live and how wrong it may be. It survives replacing the event bus, the cache, the SDK and the store, because none of those changes the bound.
LessonIf a consumer may be briefly wrong, put the state in the consumer and publish how wrong it is allowed to be. If it may never be wrong, do not pretend a cache is an optimisation.
Shown on views02 13 15
ADR-08

Fail closed by default, with fail-open as a declared per-purpose posture

Accepted

When the decision path cannot answer, does processing stop or continue?

Context
Universal fail-closed is the correct answer in a privacy design and makes this platform's availability the hard ceiling on every product that depends on it: a bad deploy here stops recommendations, personalised search, saved payment methods and half the home page. Universal fail-open is indefensible — a brief outage becomes a brief period of routine unlawful processing, and nobody will notice. A declared exception set is the only option that survives contact with the business, and it has a sharp edge of its own: it puts a legal judgement into a configuration field, where it can be changed by anyone with deploy access unless something stops them.
Decision
The default posture is fail-closed. Each purpose may declare a fail-open posture in the registry, which requires counsel approval and two-person sign-off, is visible in the records of processing, and is audited every time it is relied upon. The expectation is that the set stays small and is dominated by purposes whose lawful basis is not consent.
How it is realised on AWS
The posture is a field on the purpose version in Aurora Global, carried in the signed bundle and evaluated by the SDK when the cache is past its staleness ceiling and the regional API is unreachable. Relying on a fail-open posture emits a distinct audited event with the purpose, the enforcement point and the duration, so the exception's real-world usage is a reported number rather than an assumption.
Options weighed
  • ChosenFail closed by default; per-purpose fail-open under approval and audit: Survives the business conversation; concentrates the danger in one field and guards that field heavily.
  • RejectedUniversal fail-closed: Architecturally cleanest and makes this platform the availability ceiling for the whole product estate.
  • RejectedUniversal fail-open: Trades a brief outage for a brief period of unlawful processing that nobody will detect.
  • Right elsewhereFail to the last-known verdict indefinitely: Reasonable where withdrawals are rare and consequences mild; here it means a withdrawal can be ignored for the duration of an outage.
Consequences
What it buys
  • The dangerous choice is explicit, owned by named approvers, and visible in the same document a regulator reads.
  • Relying on fail-open is itself measured, so the exception cannot quietly become the normal path.
  • Products whose purposes are consent-based degrade safely without anyone making a judgement call in an incident.
What it costs
  • The posture field is the most dangerous single field in the system, and its guard is process rather than architecture.
  • A fail-closed default means this platform's availability is felt directly by users during an incident, which creates pressure to widen the exception set.
  • Teams will discover the posture field exists and will ask for it, each with a plausible case.
Choose differently when
If the decision path's measured availability reached a level where closed-failure was effectively never exercised, the exception set could be removed entirely — which is the right end state. If the exception set grows past a handful of purposes, the design has failed and the honest response is to attack the availability rather than widen the list.
Why it holds up over time
Every system that gates another system's behaviour faces this question, and the shape of the answer — safe default, explicit audited exceptions, measured reliance — generalises well beyond privacy.
LessonPut the unsafe option behind a named approval and then measure how often it is used. An exception you cannot count is an exception that becomes the rule.
Shown on views15 22
ADR-09

UNKNOWN is a distinct verdict, and a caller may never read it as ALLOW

Accepted

Should "there is no grant" and "I could not read whether there is a grant" be the same answer?

Context
Collapsing them into DENY is tempting and nearly right: both block processing, so why distinguish? Because they require opposite responses from the platform. No grant is the system working. An unreadable record is the system broken, and if it presents as DENY then a failing projection looks like a population of users who all happened to opt out, which is a failure that can run for days while the dashboards look plausible. The opposite collapse — into ALLOW — is the single worst thing this platform could do, and is also the one a tired client developer will reach for when the field is missing from a response.
Decision
Three verdicts exist: ALLOW, DENY and UNKNOWN. DENY means no permission; UNKNOWN means the subject's state could not be determined. A caller must handle UNKNOWN explicitly, and treating it as ALLOW is a contract violation. UNKNOWN is never returned where a purpose requires consent and the record was read successfully with no grant present — that case is DENY.
How it is realised on AWS
The decision response is a required enum rather than a boolean, with no default, so a client that ignores the field fails to deserialise rather than defaulting to permissive. The SDK ships a contract test suite that fails a build if UNKNOWN is mapped to permitted. UNKNOWN rates are reported per enforcement point, since a rising rate is a platform defect, not user behaviour.
Options weighed
  • ChosenThree verdicts, UNKNOWN explicit and non-defaultable: Distinguishes working from broken; costs every caller an extra branch and a contract test.
  • RejectedTwo verdicts, unreadable state returns DENY: Safe for the subject and makes a broken read path indistinguishable from mass opt-out for as long as it lasts.
  • RejectedBoolean permitted flag: Smallest contract and the one most likely to be defaulted to true somewhere in forty client integrations.
  • Right elsewhereVerdict plus a confidence score: Useful where enforcement is probabilistic; here a legal decision with a confidence interval is not actionable by any caller.
Consequences
What it buys
  • A failing projection or an unreachable region shows up as an UNKNOWN rate rather than as an inexplicable collapse in feature reach.
  • The permissive failure mode has to be written deliberately; it cannot be reached by omitting a field.
  • Each verdict maps to a distinct action: serve, suppress, or apply the purpose's posture.
What it costs
  • Three verdicts plus a caller error is four paths in forty integrations, and the SDK carries the burden of making that bearable.
  • An UNKNOWN still has to resolve to behaviour, which is why the posture record exists; UNKNOWN alone is not an answer to the product.
  • A contract test can be disabled, so the guarantee is social as much as technical.
Choose differently when
If enforcement moved entirely into a platform-owned access layer, the verdict contract becomes internal and the risk of permissive defaulting largely disappears. Nothing else would justify collapsing the enum.
Why it holds up over time
"Absent" and "unavailable" are different facts in every distributed system ever built, and conflating them has caused the same class of incident for forty years.
LessonMake the unsafe interpretation impossible to reach by accident. A boolean with a default is a decision someone else will make for you, badly, at 3am.
Shown on views15 22

Erasure as a protocolWhy deletion is a distributed protocol with verification rather than a command with a return code.

ADR-10

Erasure is a four-state verified protocol per target, never a boolean

Accepted

When a hundred systems have been told to delete a subject, what does the case need to record about each one?

Context
The natural model is a boolean per target: told, done. It produces a case that reaches 100% and closes, and it cannot distinguish a system that deleted the data from one that accepted the message, from one that claimed success without looking, from one that went quiet. Those are four different situations with four different responses, and a privacy programme that cannot tell them apart is confidently wrong about its own completeness — which is the most common way this capability fails, and it fails silently.
Decision
Every target of every erasure case carries four distinct states: instructed, acknowledged, attested complete, and verified by the platform. They are never collapsed. A case reports `partial` until every target reaches a terminal state, and terminal includes a recorded refusal with a reason. Reaching attested is not completion; verification is a separate act performed by the platform.
How it is realised on AWS
A Step Functions state machine per case fans out through target adapters, writing a target_outcome item per (case, target) in DynamoDB with the four-valued state, timestamps and any refusal reason. Adapters are idempotent and retried with backoff. The verification prober is a separate scheduled Fargate task reading the attested set and probing a sample, writing findings back onto the case and to the compliance dashboard.
Options weighed
  • ChosenFour states per target, verification as a platform act: Distinguishes the four real situations; costs state, adapters and a probing budget.
  • RejectedBoolean per target: Clean dashboard, and it manufactures completed erasures out of accepted messages.
  • RejectedThree states without platform verification: Trusts the attestation, which is a claim about a system by that system — exactly the evidence least worth having.
  • Right elsewhereFire-and-forget with a reconciliation sweep: Workable where every target is one owned data platform; across sixty vendors it has no per-case accountability at all.
Consequences
What it buys
  • A case's state is derived from evidence rather than from elapsed time, so "complete" means something.
  • A target that attests and still returns data becomes a named finding against a named owner.
  • Refusals are terminal and explainable, which is what lets a lawful partial fulfilment close honestly.
What it costs
  • Four states across roughly a hundred targets per case is real state and real operational surface.
  • Verification can only ever be sampled, so completeness is evidenced with a stated confidence rather than guaranteed.
  • Cases stay open longer and the dashboard looks worse than a boolean one would, which will be mistaken for the platform performing badly.
Choose differently when
If every target could expose a cryptographic proof of deletion, verification becomes deterministic and the fourth state becomes a check rather than a probe. Nothing available today does this, and a design that assumes it is assuming its hardest problem away.
Why it holds up over time
The protocol shape — instruct, acknowledge, attest, verify — is how every distributed obligation with an untrusted counterparty has to work, and it predates this domain by decades.
LessonWhen you cannot inspect the outcome, model the evidence rather than the outcome. A boolean is a claim about reality; four states are a record of what you actually know.
Shown on views05 14
ADR-11

The erasure technique is declared per store, in advance, and is part of the evidence

Accepted

Does erasure mean deletion, and if it does not always, who decides what it means for each store?

Context
A row in a transactional table can be deleted. An append-only log cannot be rewritten without destroying the integrity that makes it useful. A columnar warehouse can delete, expensively, by rewriting files. A trained model's weights cannot be edited at all. Trained-on data, immutable audit trails and backups each rule out some options, so a single definition of erasure cannot hold across the estate. Discovering this per store during a case — which is what happens when the technique is not declared in advance — means a thirty-day statutory clock spent in architectural discussion.
Decision
Each registered target declares its erasure technique in advance: hard delete, crypto-shredding of the subject's per-subject key, irreversible anonymisation, or tombstone-plus-suppression. The declared technique is registry data, is part of the case evidence, and is what the subject's confirmation describes. Where crypto-shredding is relied upon, the registry states that the ciphertext remains.
How it is realised on AWS
processing_target carries class and technique in Aurora Global. Each subject has a KMS data key per region wrapped by a regional CMK; stores relying on crypto-shredding encrypt per-subject payloads under it, and the case destroys the key version rather than the rows. The warehouse uses filter-and-re-derive driven by the suppression feed. Every case's evidence pack names the technique applied at each target.
Options weighed
  • ChosenPer-store declared technique, recorded as evidence: Honest about what is possible where; makes the awkward answers visible before a case needs them.
  • RejectedHard delete everywhere: Clean promise the estate cannot keep, which means someone will quietly not keep it.
  • RejectedCrypto-shredding everywhere: Elegant and makes every store's read path depend on a per-subject key, which is a large performance and operational cost imposed on stores that could simply delete.
  • Right elsewhereAnonymisation everywhere: Right for analytical estates where identifiability is the only concern; it fails wherever the subject's own records must be removed, not just de-identified.
Consequences
What it buys
  • An immutable log can satisfy an erasure obligation without being compromised as an audit record.
  • The subject's confirmation can describe what actually happened, store class by store class.
  • Architectural debates about what deletion means happen at registration, not inside a statutory clock.
What it costs
  • Whether destroying a key counts as erasure is jurisdiction-dependent and genuinely contested, so this design carries legal risk it cannot resolve.
  • Per-subject keys at 180 million subjects is significant key-management volume and a hard dependency on the key store's availability.
  • A technique declared years ago may no longer match how the store actually works, which makes the declaration itself something that needs re-verification.
Choose differently when
If a regulator held that crypto-shredding does not constitute erasure, every store relying on it needs a different technique and some of them have none — which would force either data-model change or a narrower retention policy upstream. Conversely, if storage engines gained first-class per-row deletion with integrity preservation, this record simplifies to hard delete almost everywhere.
Why it holds up over time
The underlying truth — that some stores cannot forget without destroying what makes them useful — is a property of append-only and derived data, not of any product, and will outlast every engine in the estate.
LessonWhen a single word in a requirement means four different operations in practice, make it four named operations in the model. Uniform language over non-uniform reality is how obligations get missed.
Shown on views11 14
ADR-12

Silence from a target is failure, not pending

Accepted

A target was instructed and has said nothing for a week. Is the case waiting, or is it broken?

Context
Treating silence as pending is comfortable and lets a case close when a timer expires, which manufactures completed erasures that never happened. Treating silence as failure means an unreliable vendor integration blocks cases, escalates to a named owner, and appears on a compliance dashboard that executives read — which generates pressure to reclassify. The asymmetry matters: a false completion is an unlawful state the organisation believes is lawful, while a false block is visible work.
Decision
A target that neither acknowledges nor attests within its declared window is treated as failed. The case stays open in `partial`, retries with backoff, escalates to that target's named owner, and appears on the compliance dashboard as a blocking gap. A case never auto-closes on a timer, and silence is never recorded as completion.
How it is realised on AWS
Each target's contractual window is registry data. The Step Functions case machine sets a timer per target, and expiry transitions that target to a blocking failure rather than to a terminal success, emitting a finding with the target's owner attached. Dashboard metrics count blocked targets by owner, and the case SLA clock is reported separately from the target state so a near-deadline case is visible before it is late.
Options weighed
  • ChosenSilence is a blocking failure with an owner: Never fabricates completion; creates visible, attributable work and the political pressure that comes with it.
  • RejectedSilence is pending; timer closes the case: A clean dashboard built on cases that closed without evidence — the failure this whole package exists to prevent.
  • RejectedSilence is pending indefinitely, no timer: Honest and useless: a case that is forever open is a case nobody acts on.
  • Right elsewhereSilence tolerated for low-risk targets by classification: Defensible where some targets hold trivially re-derivable data; here it reintroduces a judgement call per target that will drift towards leniency.
Consequences
What it buys
  • A statutory deadline is at risk visibly and in advance, rather than discovered after it passes.
  • Unreliable integrations are surfaced as an owner's problem, which is the only thing that ever gets them fixed.
  • Case state means what it says, which is what makes the subject's confirmation truthful.
What it costs
  • Blocked cases accumulate for targets with weak integrations, and the backlog is uncomfortable to look at.
  • There is standing pressure to reclassify silence as pending, and it will come with a business case each time.
  • Escalation requires every target to have a current named owner, which is an organisational dependency the platform cannot enforce.
Choose differently when
If contracts with recipients mandated machine-readable attestation with penalties, the silent case becomes rare enough that its handling matters less. Nothing about the technical design would change; the volume would.
Why it holds up over time
Treating absence of response as failure rather than success is a first principle of any protocol with an untrusted counterparty, and does not depend on the domain.
LessonDecide early which direction an unknown resolves in, and pick the direction whose failures are visible. A system that resolves unknowns optimistically is a system that lies to its owners.
Shown on views14 22
ADR-13

The suppression list is mandatory on every ingestion and restore path

Accepted

What stops an erased subject from coming back through a backup restore, a replayed event stream or a vendor re-upload?

Context
Erasure is a point-in-time act, and data flows continuously. A restore from a four-week-old backup reintroduces a subject erased three weeks ago. A replayed stream re-creates rows. A vendor's periodic file upload re-adds a contact the vendor was told to delete and did, from a copy they kept for reconciliation. Each of these is a separate re-entry path and there is no single system that owns all of them — which is precisely why a per-path solution fails: there is always one more path.
Decision
A single compact suppression list of hashed erased subject keys and erasure dates is maintained per region, and consultation is a mandatory precondition of every ingestion and restore path in the estate. A path that cannot consult it cannot be registered as a processing target. The list is retained indefinitely, because the obligation does not expire.
How it is realised on AWS
A DynamoDB table of hashed subject keys with erasure dates, replicated in-region, with a compact snapshot published to S3 and CloudFront for paths that need a local copy. The SDK exposes a suppression check alongside the decision call; the warehouse applies it as a filter on load; restore runbooks require a suppression pass before a restored dataset is made readable. Non-consultation is detected by the verification prober, which probes restored and re-ingested data specifically.
Options weighed
  • ChosenOne mandatory suppression list, consulted on every entry path: One control that holds across all re-entry paths, including the ones nobody enumerated; costs a check in the hottest write paths.
  • RejectedRe-run erasure after each restore: Correct in principle and depends on someone remembering, which is the failure mode being designed against.
  • RejectedPer-subject encryption so backups self-shred: Elegant where it applies, and it does not cover a vendor re-upload or a replayed stream carrying plaintext.
  • Right elsewhereShorter backup retention to bound the window: Genuinely effective and usually impossible: backup retention is set by recovery and legal requirements, not by privacy.
Consequences
What it buys
  • One control covers restore, replay and re-upload, including paths discovered after the design was written.
  • The list holds only hashed keys and dates, so it is cheap to replicate and carries minimal content risk.
  • Non-consultation is independently detectable, rather than relying on each path's owner to confirm it.
What it costs
  • A check on every ingestion path is a hot-path cost imposed on systems that get no benefit from it.
  • The list is pseudonymous personal data retained indefinitely, which needs its own lawful basis and its own explanation to subjects.
  • It grows forever, and the hash is only as protective as the salt management around it.
Choose differently when
If the estate's ingestion consolidated behind a single platform-owned pipeline, the check moves to one place and stops being a distributed obligation. If a jurisdiction held that retaining erased subjects' hashed keys is itself unlawful, the design would need per-subject crypto-shredding to carry the whole burden, which it cannot for plaintext re-uploads.
Why it holds up over time
Any system that erases from a continuously-fed store needs a negative index. The shape of that requirement does not change with technology.
LessonYou cannot enumerate every way data gets back in. Put the control where everything enters rather than where you believe things enter.
Shown on views11 14 22

Residency and evidenceWhere the record of permission lives, and what has to survive the request to be forgotten.

ADR-14

Residency is enforced in the data path, and consent metadata is personal data

Accepted

Is a jurisdictional boundary a property of the architecture or a column in the database?

Context
One global deployment with a residency tag per row is dramatically cheaper to run and is a single configuration error away from an unlawful transfer — and that error is undetectable from inside the system, because the data served looks correct. Isolated regional stacks make the boundary structural: there is no replica to read from, so a routing mistake fails rather than leaks. The cost is nine deployments, nine key hierarchies, no cross-border failover, and the acceptance that a region's bad day denies its own subjects' consent-based purposes. The subtler point is that the consent record is itself personal data about the subject, so the boundary applies to this platform's own stores and not only to the data it governs.
Decision
Each jurisdiction is served by its own regional stack holding the ledger, projection, identity index, case state, suppression list, keys and audit for its subjects. Nothing crosses: no replica, no key, no failover. Only the purpose registry and the signed policy bundle — which hold no personal data — replicate globally. A request carrying a subject key outside the region's jurisdiction is rejected or referred, never served from a replica that happens to hold it.
How it is realised on AWS
Regional stacks in eu-west-1/eu-central-1, us-east-1/us-west-2 and ap-south-1, each with its own DynamoDB tables, KMS CMK with no cross-region grant, Step Functions, and S3 Object Lock audit bucket. Aurora Global carries the registry. A jurisdiction router at the edge, backed by Route 53 and edge policy, directs each subject's traffic to their boundary, and the regional API rejects foreign subject keys outright rather than proxying them.
Options weighed
  • ChosenIsolated regional stacks; only definitions global: Makes the boundary structural so a routing bug fails instead of leaking; costs nine stacks and no failover.
  • RejectedOne global deployment with residency tags and policy routing: Far cheaper, and a single misconfiguration produces an undetectable unlawful transfer.
  • RejectedGlobal control plane with regional data planes, personal data regional: Essentially the chosen design with a larger global surface; rejected because any global component holding subject keys reintroduces the risk.
  • Right elsewhereSovereign operator-run deployment per jurisdiction: Right where a jurisdiction requires operational sovereignty too; here it multiplies cost without changing the data boundary.
Consequences
What it buys
  • An unlawful transfer requires a deliberate act, not a configuration mistake.
  • Adding a jurisdiction is a new deployment and new configuration, not a schema change.
  • Per-region keys mean a compromise of one boundary's key hierarchy cannot decrypt another's.
What it costs
  • A region's unavailability denies its subjects' consent-based purposes, and there is no failover to offer.
  • Nine stacks multiply operational cost, patching surface and configuration drift risk.
  • The jurisdiction router becomes the one component whose failure is an unlawful transfer, so it is now the most safety-critical routing decision in the estate.
Choose differently when
If a jurisdiction permitted consent metadata to be held outside its boundary under a recognised transfer mechanism, the stacks could consolidate and the cost falls sharply. More likely the pressure runs the other way: new jurisdictions are added, which this design absorbs and the global-with-tags design does not.
Why it holds up over time
Jurisdictional fragmentation of data protection is a decade-long trend. A design whose boundary is enforced by routing and keys scales with that trend; one whose boundary is a policy document does not survive its first audit.
LessonIf a boundary matters legally, make it impossible to cross rather than wrong to cross. A boundary enforced by correctness of configuration is a boundary you will cross by accident.
Shown on views08 16 20
ADR-15

Evidence that an erasure happened survives the erasure

Accepted

When a subject asks to be forgotten, may the platform keep the record proving it forgot them?

Context
The obligation to erase and the obligation to demonstrate compliance point in opposite directions. Honouring erasure completely would destroy the evidence that erasure occurred, leaving the organisation unable to answer the regulator's most basic question and unable to defend itself against a claim that it did nothing. Keeping full records means holding personal data about someone who explicitly asked not to be held. There is no technical resolution: the tension is real, and the only question is which way it is resolved and whether the subject is told.
Decision
The platform retains the minimum evidence needed to prove an erasure occurred — subject key, case identifier, per-target outcomes, timestamps, the technique applied — under a lawful basis of legal obligation that the subject cannot withdraw. Everything beyond that minimum is erased with the rest. The retention and its basis are stated plainly to the subject in the erasure confirmation, not buried in a policy.
How it is realised on AWS
Audit records land in S3 with Object Lock in compliance mode, hash-chained so a gap is detectable, with independent credentials and a separate blast radius from the operational stores. The erasure case removes the subject's consent entries' content while retaining the case skeleton and outcomes. The retained minimum is enumerated in the records of processing and reproduced verbatim in the subject's confirmation.
Options weighed
  • ChosenMinimum evidence retained under legal obligation, stated to the subject: Defensible and honest; it still means holding data about someone who asked not to be held.
  • RejectedErase everything including the evidence: Maximally faithful to the request and leaves the organisation unable to prove it complied, including against a false claim that it did not.
  • RejectedRetain full case detail indefinitely: Convenient for audit and keeps far more than proving compliance requires, which is the definition of excessive processing.
  • Right elsewhereHand the evidence to the subject and keep nothing: Appealing and unworkable: the organisation's obligation to demonstrate compliance cannot be discharged by data only the subject holds.
Consequences
What it buys
  • The regulator's question is answerable for erased subjects, which is otherwise the one population about whom nothing can be said.
  • The retained set is enumerable and minimal, so the trade is reviewable rather than open-ended.
  • Hash-chained write-once storage means a missing record is detectable, which is what makes the evidence worth anything.
What it costs
  • The platform holds personal data about people who asked not to be held, under a basis they cannot withdraw. That is the honest cost and it cannot be designed away.
  • The subject's confirmation must explain this, which makes a good-news message partly bad news.
  • Seven-year write-once retention of audit records is a cost and a disclosure obligation of its own.
Choose differently when
If a regulator specified a shorter or narrower evidentiary minimum, the retained set shrinks accordingly — the design already enumerates it, so this is configuration. If demonstrating compliance were accepted on aggregate rather than per-subject evidence, the retained set could become statistical and the tension largely disappears.
Why it holds up over time
The conflict between a right to erasure and a duty to demonstrate compliance is structural to every accountability regime, and will be present in whatever replaces the current ones.
LessonWhen two obligations genuinely conflict, resolve it explicitly, minimise what you keep, and tell the person. A conflict resolved silently is one you will be asked about under worse conditions.
Shown on views06 12 22

Every package used, in one table

Terms used precisely in this package, including several that are used loosely elsewhere. Where a looser alternative exists, it is named along with what it costs.

PackageWhat it isWhat it does hereConsidered instead
Purpose A declared, versioned reason for processing personal data, carrying data categories, retention, permitted systems and recipients. The unit everything else is keyed on: consent, decisions, erasure fan-out and cost all hang off it. "Use case" or "processing activity", which blur the versioned legal object with the product feature that happens to rely on it.
Lawful basis The legal justification for processing under a purpose, declared per purpose version per jurisdiction. Decides the default: no processing before a grant under consent, processing until objection under legitimate interest. Treating consent as the universal basis, which forfeits lawful processing and trains users to click through everything.
Widening change A purpose change adding a data category, a recipient, or a longer retention period. Triggers a new purpose version and fresh permission where the basis is consent; a narrowing change applies immediately. "Update", which lets a grant given for something small silently cover something larger.
Propagation ceiling The published maximum time between a withdrawal being committed and every enforcement point reflecting it — 15 minutes here. The real compliance bound. Within it, staleness is accepted; beyond it, the platform is in an incident. A latency target, which implies a performance concern rather than an accruing legal exposure.
Staleness ceiling The maximum age a local cached consent entry may reach before the purpose's failure posture applies. Converts a decision-plane outage into a bounded, declared degradation instead of silent divergence. A TTL, which is the same mechanism without a stated obligation attached to it.
Failure posture A per-purpose registry field deciding whether an unanswerable decision fails closed or open. Makes the dangerous choice explicit, approved by two people, visible in the records of processing, and audited whenever relied upon. A global fallback, which either makes this platform the availability ceiling for the estate or legalises an outage's worth of processing.
Crypto-shredding Satisfying erasure by destroying the subject's encryption key, leaving unreadable ciphertext. The only erasure technique available for append-only and integrity-protected stores. Calling it deletion, which overstates what happened and may not satisfy every jurisdiction.
Attested versus verified Attested is the target's claim that it deleted; verified is the platform's own evidence that the data is gone. Keeping them distinct is what stops a case closing on a claim by the party with the least incentive to check. "Confirmed", which collapses the two and makes every completeness statement unfalsifiable.
Suppression list The set of hashed erased subject keys consulted by every ingestion and restore path. The single control preventing resurrection by backup restore, stream replay or vendor re-upload. Re-running erasure after each restore, which depends on somebody remembering every path.
Jurisdictional boundary A set of regions serving one jurisdiction's subjects, with no replica, key or failover crossing it. Makes residency a structural property rather than a configuration value that can be set wrongly. A residency column, which is a single misconfiguration away from an undetectable unlawful transfer.
Subject key The pseudonymous internal identifier for a data subject, to which consent, cases and evidence are keyed. Lets the platform govern a person without holding their identifiers or their data. Using an email address or user id, which spreads a directly identifying value through every store here.
UNKNOWN A decision verdict meaning the subject's permission state could not be determined. Separates the system working (DENY, no grant) from the system broken, which otherwise look identical in the metrics. Folding it into DENY, which makes a failing projection indistinguishable from a population that all opted out.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.