# Architecture Decision Record

*Consent & Privacy Service · Solution Architecture v1.0 · Amazon Web Services · Security & Identity Architecture · 2026-10*

The argument these decisions serve is summarised in the [Architecture One-Pager](architecture-one-pager).

Fifteen decisions make up this architecture. Everything else across the twenty-two views is a consequence of one of them, and each record carries the question that forced it, how it is realised on AWS, the alternatives including the ones that are right for a different organisation, and the conditions under which the choice should be reversed.

> **Status of this document.** This is a design, not a report on a running system. Every rate, latency, volume, retention and cost figure in this package is a stated assumption, chosen to be defensible for a consumer SaaS of roughly 180 million registered subjects across 40 internal processing systems, 60 external recipients, 9 regional deployments and 3 jurisdictional boundaries, carrying 120,000 permission decisions a second at steady state and 25,000 rights requests a day. They are written as numbers so that they can be argued with and corrected, which vagueness does not allow.

## How to read a record

- **Question:** The forcing question: why a decision was needed at all.
- **Context:** The requirement, the scale and the constraint that make it hard.
- **Decision:** What this architecture does, stated so it can be checked.
- **How it is realised on AWS:** The concrete mechanism: which service or package, configured how, in which subscription.
- **Options weighed:** Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- **Consequences:** What the choice buys and what it costs, both kept visible.
- **Choose differently when:** The conditions that would flip the decision for your system.
- **Why it holds up over time:** What keeps the decision right as scale, staff and technology change.
- **Lesson:** The principle that transfers beyond this platform.

## Decision map

**The boundary**: What this platform owns, and what it deliberately refuses to own.

- ADR-01 · Permission is the system of record; personal data stays with its owner
- ADR-02 · The identifier graph is centralised, and treated as the most dangerous store in the design

**Permission as data**: How a permission is recorded, versioned, and made provable eighteen months later.

- ADR-03 · The consent ledger is the system of record; current state is a disposable projection
- ADR-04 · Lawful basis is declared per purpose version per jurisdiction, and widening re-asks
- ADR-05 · There is no default purpose, and an unregistered purpose is rejected rather than denied

**Enforcement**: Where the purpose question is asked, and what the answer is when it cannot be answered.

- ADR-06 · Purpose limitation is enforced at the point of use, not at collection
- ADR-07 · Consent state is pushed to local caches with a published staleness ceiling
- ADR-08 · Fail closed by default, with fail-open as a declared per-purpose posture
- ADR-09 · UNKNOWN is a distinct verdict, and a caller may never read it as ALLOW

**Erasure as a protocol**: Why deletion is a distributed protocol with verification rather than a command with a return code.

- ADR-10 · Erasure is a four-state verified protocol per target, never a boolean
- ADR-11 · The erasure technique is declared per store, in advance, and is part of the evidence
- ADR-12 · Silence from a target is failure, not pending
- ADR-13 · The suppression list is mandatory on every ingestion and restore path

**Residency and evidence**: Where the record of permission lives, and what has to survive the request to be forgotten.

- ADR-14 · Residency is enforced in the data path, and consent metadata is personal data
- ADR-15 · Evidence that an erasure happened survives the erasure

## Technology by capability

Every capability below is satisfied by an AWS service or an open-source component chosen for a property the requirement names, not for familiarity. Where a simpler or more obvious option was rejected, the reason is the architecture rather than the service.

| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| Consent ledger (system of record) | DynamoDB, keyed on subject_key with captured_at sort key, conditional writes forbidding overwrite, Streams enabled | Amazon Web Services | Aurora PostgreSQL append-only table; Kinesis as the log | A subject key is naturally high-cardinality, so partitioning by it removes hot keys entirely at 2.5 billion current-state records, and Streams gives the projection builder a durable ordered feed without a second system. | ADR-03 |
| Current-state projection | DynamoDB single-item-per-(subject, purpose, jurisdiction), version-stamped, rebuilt from the ledger | Amazon Web Services | ElastiCache as the primary store; a relational materialised view | The read is a point lookup at 120,000/second with no range scan, and the store must be rebuildable rather than durable — which rules out anything whose contents cannot be discarded. | ADR-03 |
| Purpose and lawful-basis registry | Aurora Global Database, strongly consistent in the primary region, globally readable | Amazon Web Services | DynamoDB global tables; a Git repository as the sole source | The registry is small, highly relational (purpose, version, basis, target), and needs real constraints and transactions; it holds no personal data, which is the only reason it is allowed to be global at all. | ADR-04 |
| Policy distribution to regions and SDKs | Signed bundle in S3 fronted by CloudFront, versioned, verified on load | Amazon Web Services | Direct reads from the registry; AppConfig | A decision plane must keep working with the global registry unreachable, so it needs a local artefact with a version it can report — and a signature so a cached bundle cannot be substituted. | ADR-07 |
| Withdrawal fan-out | EventBridge, one bus per region, per-subject ordering, one subscription per enforcement point | Amazon Web Services | SNS fan-out; Kafka; polling the projection | Per-enforcement-point subscriptions are what make per-point propagation lag measurable, which is the metric the whole 15-minute ceiling claim depends on. | ADR-07 |
| Rights-case orchestration | Step Functions standard workflows, one execution per case, DynamoDB for target outcomes | Amazon Web Services | A homegrown saga on SQS; Temporal | A case runs for up to thirty days across a hundred targets with retries, timers, escalations and human refusal steps; durable resumable execution with per-step history is the requirement, and the history doubles as evidence. | ADR-10 |
| Identity resolution index | DynamoDB table with per-item encryption under per-subject KMS data keys, separate authorisation, per-record read audit | Amazon Web Services | A graph database; resolution assembled per case from source systems | A graph engine's traversal power is not needed for alias lookup and would widen what a compromise yields; per-subject keys make the linkage disappear when the subject is crypto-shredded. | ADR-02 |
| Crypto-shredding | KMS data key per subject per region, wrapped by a regional CMK with no cross-region grant | Amazon Web Services | Application-managed keys in a secrets store; envelope keys per tenant | Erasure from append-only stores has to be key destruction, and the key boundary has to coincide exactly with the residency boundary, which a regional CMK with no grant enforces structurally. | ADR-11 |
| Tamper-evident evidence store | S3 with Object Lock in compliance mode, hash-chained records, independent credentials | Amazon Web Services | QLDB; an append-only table in the operational database | The evidence must survive compromise of the operational plane, so it needs a separate credential domain and a retention mode that the platform's own administrators cannot shorten. | ADR-15 |
| Batch and campaign-scale evaluation | Immutable projection snapshot exported to S3, queried with Athena | Amazon Web Services | Looping the per-request decision API; a read replica | A million-subject eligibility question must not share capacity with the read path, and a snapshot gives a single consistent version to report alongside the result. | ADR-06 |
| Compute for capture, decision and adapters | ECS on Fargate across three AZs per region, with the decision SDK in-process in consuming services | Amazon Web Services | Lambda; EKS | Steady high-throughput services with warm caches are a poor fit for per-invocation isolation, and EKS adds a control plane this platform does not need for a dozen services per region. | ADR-01 |
| Subject authentication and step-up | Cognito for subject identity with step-up for export and erasure; workforce IdP federation for agents | Amazon Web Services | The product's own session as sufficient proof | An unverified erasure request is an attack on the subject, so the verification strength has to be a property of the request type rather than of the session that happens to exist. | ADR-02 |
| Suppression list distribution | DynamoDB per region plus a compact snapshot on S3 and CloudFront for local copies | Amazon Web Services | A shared database table; a published file per consumer | Every ingestion and restore path must be able to consult it, including batch paths with no network access to the platform's APIs — so it needs both an online and an offline form. | ADR-13 |
| Observability | CloudWatch metrics dimensioned per enforcement point and per purpose, X-Ray on the decision path, findings routed to owning teams | Amazon Web Services | Platform-level aggregate dashboards | Propagation lag and UNKNOWN rate are only meaningful per enforcement point; a platform average hides the single cache that stopped consuming, which is the failure that matters. | ADR-12 |

## The decisions, and the alternatives that lost

### The boundary

*What this platform owns, and what it deliberately refuses to own.*

#### ADR-01 · Permission is the system of record; personal data stays with its owner

**Status:** Accepted  ·  **Shown on views:** 01, 02, 08

*To govern how personal data is used, should the governing platform hold that data, or only the record of what may be done with it?*

**Context.** There is an obvious design that makes everything else easy: copy the personal data into one governed store, and then consent, retention, export and erasure are all local operations on a single schema. Erasure becomes a delete you can prove by inspection. Export becomes a query. Purpose limitation becomes a column filter. The cost is that a platform whose purpose is to reduce the harm of holding data about people becomes the largest aggregation of data about people in the company, with its own residency, retention, consent and breach story for the copy, and a freshness problem for every field that moved. The alternative keeps the data where it is and makes the platform authoritative only over permission — which pushes the difficulty outward into forty systems and sixty vendors that must each be instructed, and must each be believed or verified.

**Decision.** The platform is the system of record for permission and never for personal data. It holds purposes, lawful bases, consent decisions, identifier linkage, case state and evidence. Personal data remains in the systems that own it, and those systems must ask before use. Every capability the platform offers is therefore either a decision about permission or an orchestration across systems it does not control.

**How it is realised on AWS.** The regional plane holds the consent ledger and current-state projection in DynamoDB, the identity index in a separately authorised DynamoDB table with per-subject KMS keys, and rights-case state in Step Functions plus DynamoDB. No product data is replicated in. Consuming systems integrate three ways: an in-process SDK that calls the decision API, an EventBridge subscription for withdrawals, and an adapter for erasure instructions. The purpose registry in Aurora Global holds the authoritative list of which systems and recipients each purpose touches.

| Option | Verdict | Reasoning |
|---|---|---|
| Permission as system of record; data federated with its owner | Chosen | Keeps the blast radius of a compromise to permission metadata; pays with orchestration, verification and an incomplete-map risk. |
| Central governed copy of personal data | Rejected | Provable single-point erasure, and it builds the thing the platform exists to limit, with freshness and residency problems for every copied field. |
| Pure policy-decision point with no state at all | Rejected | Cannot answer what was permitted last March, cannot run a rights case, and cannot prove anything to a regulator. |
| Embedded governance in each system, no central plane | Right elsewhere | Right for a small estate with one data store; at forty systems it guarantees forty inconsistent interpretations of the same consent. |

**What it buys**

- A compromise of this platform yields what people refused, not who they are and what they did — a materially smaller harm than the central-copy alternative.
- Each owning system keeps responsibility for its own data, which is where the knowledge of how to delete it correctly actually lives.
- The platform's own residency, retention and consent story stays small enough to reason about.

**What it costs**

- Erasure can never be proved by inspection. It is instructed, attested and sampled, and the residual gap is a measured quantity rather than zero.
- The purpose-to-system map becomes a correctness dependency: a system missing from it is a system that will not be told.
- Purpose limitation must be enforced on other teams' read paths, which makes adoption a product problem as much as an architectural one.

**Choose differently when.** If the estate consolidated onto one or two data platforms with a shared access layer, the central-copy design's costs mostly disappear and its provability becomes worth having. The decision should also be revisited if regulators begin to require demonstrated rather than evidenced completeness, because no federated design can supply that.

**Why it holds up over time.** This is a statement about what kind of thing the platform is authoritative over. It survives replacing every service named here, every regulation cited, and the entire rights catalogue, because none of those changes whether permission and data are the same object.

> **Lesson.** Decide what your platform is the truth about before deciding what it stores. A governance system that centralises the thing it governs usually inverts its own goal.

#### ADR-02 · The identifier graph is centralised, and treated as the most dangerous store in the design

**Status:** Accepted  ·  **Shown on views:** 12, 20, 21

*To reach every copy of a person, something must know every identifier they are known by. Should that knowledge be centralised, federated, or assembled per request and discarded?*

**Context.** A subject is a user id in one system, a device identifier in another, a hashed email at a vendor, a customer number in billing, and three anonymous session identifiers that were later linked. An erasure that reaches one of those is a failure that looks exactly like a success. Completeness therefore requires a resolution step, and the question is where the linkage lives. Central resolution is the only option that can demonstrate completeness, and it builds a store that maps every pseudonym to every person — the single most valuable target in the company, created by the team whose job is reducing privacy risk. Federated resolution avoids that store and makes completeness unprovable: each system answers for itself, and nobody can say whether the set was whole.

**Decision.** The identifier graph is centralised, per region, as a first-class store — and is then treated as the most constrained component in the architecture. It is authorised separately from the decision path, encrypted with per-subject keys, queryable only in the context of an open rights case or a decision evaluation, and audited at individual-record granularity. Holding it is accepted as the design's largest residual risk, stated rather than mitigated away.

**How it is realised on AWS.** A dedicated DynamoDB table holds identifier hashes mapped to subject keys, with items encrypted under the subject's own KMS data key so that destroying that key removes the linkage as well as the data. Access requires a case-bound role issued for a specific case id with a 30-minute expiry; the resolver service is the only principal with read access and emits one audit record per item returned to S3 Object Lock. Decision-path credentials cannot read the table at all.

| Option | Verdict | Reasoning |
|---|---|---|
| Central graph, maximally constrained and audited | Chosen | The only design that can evidence completeness; concentrates risk in one store and says so. |
| Federated resolution, each system answers for itself | Rejected | No honeypot, and no way to state whether an erasure reached everything — which makes every completion claim unfalsifiable. |
| Assembled per case from live systems, then discarded | Rejected | Attractive until a case must be reopened or audited, at which point the evidence of what was resolved no longer exists. |
| No resolution; subject must name their own identifiers | Right elsewhere | Defensible for a single-product company with one identifier; at this scale it pushes an impossible task onto the person least equipped for it. |

**What it buys**

- An erasure case can state which identifiers it covered, which makes a reopened case and an audit both tractable.
- Per-subject encryption means crypto-shredding removes the linkage, so the graph does not outlive the subjects in it.
- Per-record read auditing is the only evidence that the graph was not browsed, and it exists from day one rather than after an incident.

**What it costs**

- The company now holds a store whose compromise is worse than most of the breaches the platform was built to limit.
- Per-record audit at case volume is a real cost that will be questioned in the first budget review.
- Linking identifiers is itself processing, and needs its own lawful basis and its own entry in the records of processing.

**Choose differently when.** If the estate adopted a single pseudonymous subject identifier end to end, the graph shrinks to a trivial mapping and most of this record's cost disappears. Conversely, if a jurisdiction holds that maintaining such a graph is unlawful irrespective of purpose, resolution must become federated and the completeness claim must be weakened in public.

**Why it holds up over time.** The need to resolve a person to their aliases is intrinsic to acting on their behalf across systems, and no technology removes it. What may change is whether that resolution is allowed to be stored, which is the one condition that would force a redesign.

> **Lesson.** When a privacy control requires building a privacy risk, say so in the architecture rather than in a footnote. A risk that is named gets constrained; a risk that is implied gets reused.

### Permission as data

*How a permission is recorded, versioned, and made provable eighteen months later.*

#### ADR-03 · The consent ledger is the system of record; current state is a disposable projection

**Status:** Accepted  ·  **Shown on views:** 11, 12, 13

*Should a person's current permission be stored as a row that is updated when they change their mind, or derived from an append-only record of every decision they ever made?*

**Context.** A consent table is the natural implementation and it loses the question everybody eventually asks. "Were we allowed to send that email on 14 March" cannot be answered from a row that was overwritten in April, and a backup restore answers it with whatever happened to be in last night's snapshot. An append-only ledger answers it exactly, at the cost of a derived read path, a rebuild story, and an eventual-consistency window between a capture and the state every reader sees.

**Decision.** Every permission decision is appended to a per-region, per-subject-partitioned ledger and is never updated or deleted inside its retention period. Corrections are compensating entries. The current-state projection is derived, rebuildable, and carries the ledger version it was built from; every decision the platform returns names that version. RPO 0 for the ledger, RPO 60 s for the projection.

**How it is realised on AWS.** The ledger is a DynamoDB table keyed on subject_key with a sort key of captured_at, written with a condition that forbids overwrite, and acknowledged only after a regional quorum commit. DynamoDB Streams feeds a Fargate projection builder that maintains the current-state table and emits withdrawal events to EventBridge. A full projection rebuild runs from the stream's retained history and the ledger partition scan; archive beyond the hot window lands in S3 and is replayable.

| Option | Verdict | Reasoning |
|---|---|---|
| Append-only ledger as SoR, state as projection | Chosen | Point-in-time truth, free audit trail, rebuildable reads; pays with a derived read path and a staleness window. |
| Mutable consent table with an audit trigger | Rejected | Simpler and immediately consistent, and the audit trail is a side-effect nobody tests until it is the only evidence. |
| Event sourcing with no materialised state at all | Rejected | Correct and far too slow for a 120,000/second read path; the projection is not optional at this scale. |
| Bitemporal relational model | Right elsewhere | Right where the whole estate is already relational and volumes are modest; at billions of current-state records the write amplification is unattractive. |

**What it buys**

- Point-in-time reproduction is a query rather than a reconstruction, which is what removes the trough from the regulator journey.
- A bad projection deploy, a corrupt partition or a wrong derivation rule is repaired by rebuilding, not by surgery on live consent.
- A grant after a withdrawal is two facts, so a dispute about sequence is answerable.

**What it costs**

- Readers see state that may be up to 60 seconds behind a capture, which is why the decision response carries a version.
- The 60-minute projection RTO rests entirely on achieving ≥ 20× real-time replay, which is an assumption and not a measurement.
- Ledger growth is unbounded by design — 40 TB over seven years on the stated assumptions — and needs tiering rather than pruning.

**Choose differently when.** If retention requirements collapsed to months rather than years and nobody needed point-in-time answers, a mutable table with change data capture would be cheaper and sufficient. The decision should also be revisited if replay cannot be made fast enough to hold a credible RTO, because then the projection has quietly become the system of record whatever the diagram says.

**Why it holds up over time.** "The record of decisions is the truth and the current state is a view of it" is a modelling choice that survives every storage technology. It is the same decision an accounting ledger made, for the same reason.

> **Lesson.** If anyone will ever ask what the state was, store the decisions and derive the state. Retrofitting history onto a mutable table is not possible, only approximated.

#### ADR-04 · Lawful basis is declared per purpose version per jurisdiction, and widening re-asks

**Status:** Accepted  ·  **Shown on views:** 12, 18, 19

*Is a purpose one thing with one legal justification, or a versioned thing whose justification depends on where the subject lives?*

**Context.** The same processing — recommending content from behaviour — is consent in one jurisdiction, legitimate interest in another, and arguably contractual necessity in a third. Modelling basis as a property of the purpose forces one of three bad outcomes: the strictest basis applied everywhere, which forfeits lawful processing; the loosest applied everywhere, which is unlawful; or a purpose duplicated per jurisdiction, which makes a withdrawal in one place silently fail to cover the same processing described under another name. Separately, purposes change. Adding a data category or a recipient to an existing purpose is a different thing from narrowing it, and treating both as an edit means a grant given for something small silently covers something larger.

**Decision.** Lawful basis is an attribute of the triple (purpose, version, jurisdiction) and exactly one basis is declared per triple. Every purpose change produces a new version. A change that widens scope — a new data category, a new recipient, a longer retention period — is flagged as widening and requires fresh permission from subjects whose basis is consent; a narrowing change applies immediately to existing permissions and does not re-ask.

**How it is realised on AWS.** purpose, purpose_version and lawful_basis are three tables in Aurora Global, with basis keyed on (purpose_id, version, jurisdiction) and carrying the notice version to show. The widening flag is set in reviewed policy-as-code and cannot be set by the proposing team alone; the release path gates on it. Regional decision planes evaluate basis from the signed bundle, so a jurisdiction's rule is applied by the plane that serves that jurisdiction.

| Option | Verdict | Reasoning |
|---|---|---|
| Basis per (purpose, version, jurisdiction); widening re-asks | Chosen | Expresses reality; makes scope creep visible as a flag with an approval, rather than invisible as an edit. |
| One basis per purpose, strictest wins globally | Rejected | Safe, simple, and forfeits lawful processing the organisation is entitled to — which guarantees the platform is bypassed. |
| A separate purpose per jurisdiction | Rejected | Looks tidy until a withdrawal covers one purpose and not its twin, which is the failure mode hardest to detect. |
| Basis decided at evaluation time from a rules engine | Right elsewhere | Right where jurisdictions change faster than releases; here it makes the basis unprovable after the fact, which defeats the evidence requirement. |

**What it buys**

- A decision response can name the basis it relied on, which is the first thing a regulator asks after the verdict.
- Scope creep becomes a visible flag with two approvers rather than an unremarkable configuration edit.
- A new jurisdiction is a new basis row and a new regional deployment, not a schema change.

**What it costs**

- Re-asking is unpopular, and the pressure to classify a widening change as narrowing will be constant and will come from people with good intentions.
- An expiry or re-ask wave concentrates prompts and looks to subjects like the banner returning, which needs operational spreading the model does not provide.
- Purpose version proliferation makes the registry harder to read than a single row per purpose would be.

**Choose differently when.** If the organisation operated in a single jurisdiction, basis collapses to a purpose attribute and the versioning can be simplified. If regulators converged on a common basis taxonomy and a common test for what counts as widening, the flag could become derived rather than declared.

**Why it holds up over time.** Jurisdictional divergence in data protection is increasing, not decreasing. A model that treats basis as varying by place and time will fit the next decade better than one that treats it as a property of the processing.

> **Lesson.** Model the dimension your obligations actually vary along, even when today's estate has only one value in it. Adding a dimension later means rewriting every record you already relied on.

#### ADR-05 · There is no default purpose, and an unregistered purpose is rejected rather than denied

**Status:** Accepted  ·  **Shown on views:** 15, 19

*When a caller asks about a purpose the registry has never heard of, is the right answer DENY or an error?*

**Context.** DENY is the tempting answer because it is safe: nothing is permitted, no data is used, and the caller handles it like any other refusal. It is also the answer that hides the only interesting fact in the situation — that something in the estate is processing data under a name nobody declared, which is either a typo that will silently suppress lawful processing or a genuine undeclared purpose, and those need opposite responses. A registry with a permissive default is worse again: it turns the absence of a declaration into permission.

**Decision.** The registry has no default purpose and refuses to register one that lacks a description, data categories or a retention period. A decision call naming an unregistered purpose returns a typed caller error, not a verdict, and raises an unregistered-use finding against the calling system. Processing cannot proceed, but the caller is told it asked a bad question rather than that it received a lawful refusal.

**How it is realised on AWS.** The decision API validates the purpose identifier against the signed bundle before evaluating anything and returns a 4xx with a distinct error code. The completeness gate in the release path rejects an incomplete purpose definition at review time. Unregistered-use findings are emitted to CloudWatch with the calling workload identity attached and routed to that system's owner.

| Option | Verdict | Reasoning |
|---|---|---|
| Typed rejection plus a finding | Chosen | Surfaces the real problem and still blocks processing; costs callers one more error path to handle. |
| Return DENY | Rejected | Safe and silent. A typo becomes an unexplained drop in a feature's reach, and an undeclared purpose becomes invisible. |
| Permissive default for unregistered purposes | Rejected | Turns an omission into a permission; the opposite of what the platform is for. |
| Auto-register on first use, pending review | Right elsewhere | Attractive for a fast-moving startup wanting adoption; here it means production processing under an unapproved basis, which is the thing being prevented. |

**What it buys**

- An undeclared purpose is discovered from traffic rather than from an audit, which is the weakest detection in the package and this is most of it.
- A typo fails loudly in integration testing instead of quietly in production.
- The registry stays a complete statement of what the company does, which is what makes records-of-processing generation possible.

**What it costs**

- Callers must handle three verdicts and an error, which is a contract the SDK has to make easy or teams will collapse it.
- A rejected call blocks a feature, so the first production occurrence will be experienced as the platform breaking something.
- Detection depends on the system declaring the purpose it is processing under; a path nobody instrumented is still invisible.

**Choose differently when.** If the estate reached a point where purpose declaration was enforced structurally — data access impossible without a purpose token — the finding becomes redundant because the call could not be made. Until then this is the only signal.

**Why it holds up over time.** The distinction between "not permitted" and "not a question I can answer" is a general API-design truth and will outlast this domain entirely.

> **Lesson.** Never let a malformed question return a well-formed answer. Collapsing the two hides the only diagnosis worth having.

### Enforcement

*Where the purpose question is asked, and what the answer is when it cannot be answered.*

#### ADR-06 · Purpose limitation is enforced at the point of use, not at collection

**Status:** Accepted  ·  **Shown on views:** 10, 15

*Where is a purpose actually limited: by never collecting the data, by checking at every read, or by tagging rows and filtering at query time?*

**Context.** Collection-time filtering is the cheapest and most irreversible: if the data was never stored, no purpose can misuse it, and no runtime dependency exists. It also cannot honour a later grant — a subject who opts in next month gets nothing, because the history was discarded. Row tagging in the data plane is the most elegant and requires every query engine in the estate to understand and respect the tags, which is a programme rather than a design. Read-time decision is correct for every case and makes this platform a synchronous dependency of the hottest read paths in the company.

**Decision.** Purpose limitation is enforced where data is read, through a decision call on the read path, with the result carrying the ledger version it was computed from. Collection remains governed by the registry's data categories, but filtering at collection is not the enforcement mechanism. Tag-and-filter in the data plane is adopted only for the analytics warehouse, where the query engine is one system under one owner.

**How it is realised on AWS.** A thin client SDK embedded in each of the 40 consuming services evaluates from an in-process cache, falling back to the regional decision API. The call sits inside the service's own read path rather than at its edge, so a purpose question is asked where the data is actually used. The warehouse receives purpose tags plus the suppression feed and filters at query time, which is the one place in the estate where that is cheaper than asking.

| Option | Verdict | Reasoning |
|---|---|---|
| Read-time decision at the point of use | Chosen | Correct in every case including later grants; couples hot read paths to this platform's availability and latency. |
| Collection-time filtering | Rejected | Cheapest and safest, and permanently forfeits data a subject later consents to — an irreversible answer to a reversible question. |
| Row-level purpose tags filtered in the data plane | Rejected | Best end state and a multi-year programme across every engine; adopted only for the single-owner warehouse. |
| Periodic batch purge of data whose purpose lapsed | Right elsewhere | Adequate where reads are infrequent and latency-insensitive; here it leaves a window in which unlawful reads are routine. |

**What it buys**

- A later grant is honoured, because the data still exists and the decision changes.
- A withdrawal takes effect on the next read rather than on the next purge cycle.
- The enforcement point is where the context is, so a denial can be handled gracefully by the feature that knows what to do instead.

**What it costs**

- Every hot read path in the company now depends on a decision answer, which makes the next two records mandatory rather than optional.
- Forty integrations means the real enforcement behaviour is as good as the least-updated client library.
- Teams will ask to cache the verdict for longer than the staleness ceiling allows, and each request will be individually reasonable.

**Choose differently when.** If the estate consolidated reads behind one or two access layers, tag-and-filter in the data plane becomes achievable and is strictly better: no runtime dependency, no stale verdicts. If the decision call cannot be made cheap enough to survive the hottest path, read-time enforcement will be negotiated away in practice whatever the architecture says, and the honest response is to move to tagging rather than to pretend.

**Why it holds up over time.** The choice between limiting at collection and limiting at use is independent of technology and is the central design question of any data governance system. The answer here depends on whether consent is reversible, and it is.

> **Lesson.** An irreversible enforcement mechanism cannot implement a reversible policy. If the subject can change their mind, the control has to be evaluated, not applied once.

#### ADR-07 · Consent state is pushed to local caches with a published staleness ceiling

**Status:** Accepted  ·  **Shown on views:** 02, 13, 15

*Should every enforcement point hold a local copy of consent state, or ask the platform for every decision?*

**Context.** Pull gives one authoritative answer and makes this platform a synchronous dependency of 120,000 reads a second — its availability becomes the ceiling on every product's availability, and its worst day is everybody's worst day. Push gives every enforcement point a fast local answer and moves the entire risk into staleness: for some bounded period, a service will cheerfully say yes to something the subject has already refused, which is not a performance defect but unlawful processing. The hybrid that most systems land on — local cache, remote fallback — only works if the staleness is measured where it actually occurs rather than assumed to be small.

**Decision.** Consent state is pushed. Each enforcement point maintains an in-process last-known-good cache, invalidated by the withdrawal event stream and refreshed from the projection, and every verdict it returns carries the ledger version it was computed from and its own age. Staleness is measured per enforcement point, not as a platform average, and is bounded by a published fifteen-minute ceiling whose breach is an incident with a notification path rather than a line on a dashboard.

**How it is realised on AWS.** The ledger's stream drives EventBridge, which fans withdrawal events to each consuming service's subscription; the SDK drops the affected subject's cached entries on receipt and re-evaluates on next use. Each SDK instance reports its oldest unrefreshed entry as a CloudWatch metric dimensioned by enforcement point, and a breach of the ceiling pages the platform on-call and the owning team together.

| Option | Verdict | Reasoning |
|---|---|---|
| Push with per-point staleness measurement and a published ceiling | Chosen | Decouples product availability from this platform; makes staleness the one risk and makes it visible where it happens. |
| Pull a decision per read | Rejected | Always current, and it makes a consent service a hard dependency of every product read path in the company. |
| Short-TTL cache with no invalidation | Rejected | Looks like push and is not: the TTL is the propagation guarantee, so the ceiling becomes the TTL and cannot be improved by anything. |
| Central enforcement proxy holding state for callers | Right elsewhere | Right where client libraries cannot be changed; adds a hop and a tier that must be more available than the thing it fronts. |

**What it buys**

- The decision is an in-process lookup at p99 ≤ 10 ms, which keeps it out of every latency budget in the estate.
- A regional decision-plane outage becomes a bounded staleness event rather than a company-wide failure.
- Measuring per point exposes the one cache that stopped consuming three hours ago, which an average would hide completely.

**What it costs**

- There is a window in which a service says yes to something already refused, and it is a compliance exposure rather than a latency statistic.
- Every safety guard has to exist in the client as well as the platform, and the clients are owned by other teams.
- The ceiling, not the event, is the actual guarantee — a lost invalidation is corrected only when the entry ages out.

**Choose differently when.** If the organisation's read volume were low enough that a central decision service could credibly be more available than every caller, pull is simpler and strictly more correct. The decision should also be revisited if client library heterogeneity becomes unmanageable, because this design silently delegates compliance behaviour to those libraries.

**Why it holds up over time.** This is a statement about where a copy of the truth is allowed to live and how wrong it may be. It survives replacing the event bus, the cache, the SDK and the store, because none of those changes the bound.

> **Lesson.** If a consumer may be briefly wrong, put the state in the consumer and publish how wrong it is allowed to be. If it may never be wrong, do not pretend a cache is an optimisation.

#### ADR-08 · Fail closed by default, with fail-open as a declared per-purpose posture

**Status:** Accepted  ·  **Shown on views:** 15, 22

*When the decision path cannot answer, does processing stop or continue?*

**Context.** Universal fail-closed is the correct answer in a privacy design and makes this platform's availability the hard ceiling on every product that depends on it: a bad deploy here stops recommendations, personalised search, saved payment methods and half the home page. Universal fail-open is indefensible — a brief outage becomes a brief period of routine unlawful processing, and nobody will notice. A declared exception set is the only option that survives contact with the business, and it has a sharp edge of its own: it puts a legal judgement into a configuration field, where it can be changed by anyone with deploy access unless something stops them.

**Decision.** The default posture is fail-closed. Each purpose may declare a fail-open posture in the registry, which requires counsel approval and two-person sign-off, is visible in the records of processing, and is audited every time it is relied upon. The expectation is that the set stays small and is dominated by purposes whose lawful basis is not consent.

**How it is realised on AWS.** The posture is a field on the purpose version in Aurora Global, carried in the signed bundle and evaluated by the SDK when the cache is past its staleness ceiling and the regional API is unreachable. Relying on a fail-open posture emits a distinct audited event with the purpose, the enforcement point and the duration, so the exception's real-world usage is a reported number rather than an assumption.

| Option | Verdict | Reasoning |
|---|---|---|
| Fail closed by default; per-purpose fail-open under approval and audit | Chosen | Survives the business conversation; concentrates the danger in one field and guards that field heavily. |
| Universal fail-closed | Rejected | Architecturally cleanest and makes this platform the availability ceiling for the whole product estate. |
| Universal fail-open | Rejected | Trades a brief outage for a brief period of unlawful processing that nobody will detect. |
| Fail to the last-known verdict indefinitely | Right elsewhere | Reasonable where withdrawals are rare and consequences mild; here it means a withdrawal can be ignored for the duration of an outage. |

**What it buys**

- The dangerous choice is explicit, owned by named approvers, and visible in the same document a regulator reads.
- Relying on fail-open is itself measured, so the exception cannot quietly become the normal path.
- Products whose purposes are consent-based degrade safely without anyone making a judgement call in an incident.

**What it costs**

- The posture field is the most dangerous single field in the system, and its guard is process rather than architecture.
- A fail-closed default means this platform's availability is felt directly by users during an incident, which creates pressure to widen the exception set.
- Teams will discover the posture field exists and will ask for it, each with a plausible case.

**Choose differently when.** If the decision path's measured availability reached a level where closed-failure was effectively never exercised, the exception set could be removed entirely — which is the right end state. If the exception set grows past a handful of purposes, the design has failed and the honest response is to attack the availability rather than widen the list.

**Why it holds up over time.** Every system that gates another system's behaviour faces this question, and the shape of the answer — safe default, explicit audited exceptions, measured reliance — generalises well beyond privacy.

> **Lesson.** Put the unsafe option behind a named approval and then measure how often it is used. An exception you cannot count is an exception that becomes the rule.

#### ADR-09 · UNKNOWN is a distinct verdict, and a caller may never read it as ALLOW

**Status:** Accepted  ·  **Shown on views:** 15, 22

*Should "there is no grant" and "I could not read whether there is a grant" be the same answer?*

**Context.** Collapsing them into DENY is tempting and nearly right: both block processing, so why distinguish? Because they require opposite responses from the platform. No grant is the system working. An unreadable record is the system broken, and if it presents as DENY then a failing projection looks like a population of users who all happened to opt out, which is a failure that can run for days while the dashboards look plausible. The opposite collapse — into ALLOW — is the single worst thing this platform could do, and is also the one a tired client developer will reach for when the field is missing from a response.

**Decision.** Three verdicts exist: ALLOW, DENY and UNKNOWN. DENY means no permission; UNKNOWN means the subject's state could not be determined. A caller must handle UNKNOWN explicitly, and treating it as ALLOW is a contract violation. UNKNOWN is never returned where a purpose requires consent and the record was read successfully with no grant present — that case is DENY.

**How it is realised on AWS.** The decision response is a required enum rather than a boolean, with no default, so a client that ignores the field fails to deserialise rather than defaulting to permissive. The SDK ships a contract test suite that fails a build if UNKNOWN is mapped to permitted. UNKNOWN rates are reported per enforcement point, since a rising rate is a platform defect, not user behaviour.

| Option | Verdict | Reasoning |
|---|---|---|
| Three verdicts, UNKNOWN explicit and non-defaultable | Chosen | Distinguishes working from broken; costs every caller an extra branch and a contract test. |
| Two verdicts, unreadable state returns DENY | Rejected | Safe for the subject and makes a broken read path indistinguishable from mass opt-out for as long as it lasts. |
| Boolean permitted flag | Rejected | Smallest contract and the one most likely to be defaulted to true somewhere in forty client integrations. |
| Verdict plus a confidence score | Right elsewhere | Useful where enforcement is probabilistic; here a legal decision with a confidence interval is not actionable by any caller. |

**What it buys**

- A failing projection or an unreachable region shows up as an UNKNOWN rate rather than as an inexplicable collapse in feature reach.
- The permissive failure mode has to be written deliberately; it cannot be reached by omitting a field.
- Each verdict maps to a distinct action: serve, suppress, or apply the purpose's posture.

**What it costs**

- Three verdicts plus a caller error is four paths in forty integrations, and the SDK carries the burden of making that bearable.
- An UNKNOWN still has to resolve to behaviour, which is why the posture record exists; UNKNOWN alone is not an answer to the product.
- A contract test can be disabled, so the guarantee is social as much as technical.

**Choose differently when.** If enforcement moved entirely into a platform-owned access layer, the verdict contract becomes internal and the risk of permissive defaulting largely disappears. Nothing else would justify collapsing the enum.

**Why it holds up over time.** "Absent" and "unavailable" are different facts in every distributed system ever built, and conflating them has caused the same class of incident for forty years.

> **Lesson.** Make the unsafe interpretation impossible to reach by accident. A boolean with a default is a decision someone else will make for you, badly, at 3am.

### Erasure as a protocol

*Why deletion is a distributed protocol with verification rather than a command with a return code.*

#### ADR-10 · Erasure is a four-state verified protocol per target, never a boolean

**Status:** Accepted  ·  **Shown on views:** 05, 14

*When a hundred systems have been told to delete a subject, what does the case need to record about each one?*

**Context.** The natural model is a boolean per target: told, done. It produces a case that reaches 100% and closes, and it cannot distinguish a system that deleted the data from one that accepted the message, from one that claimed success without looking, from one that went quiet. Those are four different situations with four different responses, and a privacy programme that cannot tell them apart is confidently wrong about its own completeness — which is the most common way this capability fails, and it fails silently.

**Decision.** Every target of every erasure case carries four distinct states: instructed, acknowledged, attested complete, and verified by the platform. They are never collapsed. A case reports `partial` until every target reaches a terminal state, and terminal includes a recorded refusal with a reason. Reaching attested is not completion; verification is a separate act performed by the platform.

**How it is realised on AWS.** A Step Functions state machine per case fans out through target adapters, writing a target_outcome item per (case, target) in DynamoDB with the four-valued state, timestamps and any refusal reason. Adapters are idempotent and retried with backoff. The verification prober is a separate scheduled Fargate task reading the attested set and probing a sample, writing findings back onto the case and to the compliance dashboard.

| Option | Verdict | Reasoning |
|---|---|---|
| Four states per target, verification as a platform act | Chosen | Distinguishes the four real situations; costs state, adapters and a probing budget. |
| Boolean per target | Rejected | Clean dashboard, and it manufactures completed erasures out of accepted messages. |
| Three states without platform verification | Rejected | Trusts the attestation, which is a claim about a system by that system — exactly the evidence least worth having. |
| Fire-and-forget with a reconciliation sweep | Right elsewhere | Workable where every target is one owned data platform; across sixty vendors it has no per-case accountability at all. |

**What it buys**

- A case's state is derived from evidence rather than from elapsed time, so "complete" means something.
- A target that attests and still returns data becomes a named finding against a named owner.
- Refusals are terminal and explainable, which is what lets a lawful partial fulfilment close honestly.

**What it costs**

- Four states across roughly a hundred targets per case is real state and real operational surface.
- Verification can only ever be sampled, so completeness is evidenced with a stated confidence rather than guaranteed.
- Cases stay open longer and the dashboard looks worse than a boolean one would, which will be mistaken for the platform performing badly.

**Choose differently when.** If every target could expose a cryptographic proof of deletion, verification becomes deterministic and the fourth state becomes a check rather than a probe. Nothing available today does this, and a design that assumes it is assuming its hardest problem away.

**Why it holds up over time.** The protocol shape — instruct, acknowledge, attest, verify — is how every distributed obligation with an untrusted counterparty has to work, and it predates this domain by decades.

> **Lesson.** When you cannot inspect the outcome, model the evidence rather than the outcome. A boolean is a claim about reality; four states are a record of what you actually know.

#### ADR-11 · The erasure technique is declared per store, in advance, and is part of the evidence

**Status:** Accepted  ·  **Shown on views:** 11, 14

*Does erasure mean deletion, and if it does not always, who decides what it means for each store?*

**Context.** A row in a transactional table can be deleted. An append-only log cannot be rewritten without destroying the integrity that makes it useful. A columnar warehouse can delete, expensively, by rewriting files. A trained model's weights cannot be edited at all. Trained-on data, immutable audit trails and backups each rule out some options, so a single definition of erasure cannot hold across the estate. Discovering this per store during a case — which is what happens when the technique is not declared in advance — means a thirty-day statutory clock spent in architectural discussion.

**Decision.** Each registered target declares its erasure technique in advance: hard delete, crypto-shredding of the subject's per-subject key, irreversible anonymisation, or tombstone-plus-suppression. The declared technique is registry data, is part of the case evidence, and is what the subject's confirmation describes. Where crypto-shredding is relied upon, the registry states that the ciphertext remains.

**How it is realised on AWS.** processing_target carries class and technique in Aurora Global. Each subject has a KMS data key per region wrapped by a regional CMK; stores relying on crypto-shredding encrypt per-subject payloads under it, and the case destroys the key version rather than the rows. The warehouse uses filter-and-re-derive driven by the suppression feed. Every case's evidence pack names the technique applied at each target.

| Option | Verdict | Reasoning |
|---|---|---|
| Per-store declared technique, recorded as evidence | Chosen | Honest about what is possible where; makes the awkward answers visible before a case needs them. |
| Hard delete everywhere | Rejected | Clean promise the estate cannot keep, which means someone will quietly not keep it. |
| Crypto-shredding everywhere | Rejected | Elegant and makes every store's read path depend on a per-subject key, which is a large performance and operational cost imposed on stores that could simply delete. |
| Anonymisation everywhere | Right elsewhere | Right for analytical estates where identifiability is the only concern; it fails wherever the subject's own records must be removed, not just de-identified. |

**What it buys**

- An immutable log can satisfy an erasure obligation without being compromised as an audit record.
- The subject's confirmation can describe what actually happened, store class by store class.
- Architectural debates about what deletion means happen at registration, not inside a statutory clock.

**What it costs**

- Whether destroying a key counts as erasure is jurisdiction-dependent and genuinely contested, so this design carries legal risk it cannot resolve.
- Per-subject keys at 180 million subjects is significant key-management volume and a hard dependency on the key store's availability.
- A technique declared years ago may no longer match how the store actually works, which makes the declaration itself something that needs re-verification.

**Choose differently when.** If a regulator held that crypto-shredding does not constitute erasure, every store relying on it needs a different technique and some of them have none — which would force either data-model change or a narrower retention policy upstream. Conversely, if storage engines gained first-class per-row deletion with integrity preservation, this record simplifies to hard delete almost everywhere.

**Why it holds up over time.** The underlying truth — that some stores cannot forget without destroying what makes them useful — is a property of append-only and derived data, not of any product, and will outlast every engine in the estate.

> **Lesson.** When a single word in a requirement means four different operations in practice, make it four named operations in the model. Uniform language over non-uniform reality is how obligations get missed.

#### ADR-12 · Silence from a target is failure, not pending

**Status:** Accepted  ·  **Shown on views:** 14, 22

*A target was instructed and has said nothing for a week. Is the case waiting, or is it broken?*

**Context.** Treating silence as pending is comfortable and lets a case close when a timer expires, which manufactures completed erasures that never happened. Treating silence as failure means an unreliable vendor integration blocks cases, escalates to a named owner, and appears on a compliance dashboard that executives read — which generates pressure to reclassify. The asymmetry matters: a false completion is an unlawful state the organisation believes is lawful, while a false block is visible work.

**Decision.** A target that neither acknowledges nor attests within its declared window is treated as failed. The case stays open in `partial`, retries with backoff, escalates to that target's named owner, and appears on the compliance dashboard as a blocking gap. A case never auto-closes on a timer, and silence is never recorded as completion.

**How it is realised on AWS.** Each target's contractual window is registry data. The Step Functions case machine sets a timer per target, and expiry transitions that target to a blocking failure rather than to a terminal success, emitting a finding with the target's owner attached. Dashboard metrics count blocked targets by owner, and the case SLA clock is reported separately from the target state so a near-deadline case is visible before it is late.

| Option | Verdict | Reasoning |
|---|---|---|
| Silence is a blocking failure with an owner | Chosen | Never fabricates completion; creates visible, attributable work and the political pressure that comes with it. |
| Silence is pending; timer closes the case | Rejected | A clean dashboard built on cases that closed without evidence — the failure this whole package exists to prevent. |
| Silence is pending indefinitely, no timer | Rejected | Honest and useless: a case that is forever open is a case nobody acts on. |
| Silence tolerated for low-risk targets by classification | Right elsewhere | Defensible where some targets hold trivially re-derivable data; here it reintroduces a judgement call per target that will drift towards leniency. |

**What it buys**

- A statutory deadline is at risk visibly and in advance, rather than discovered after it passes.
- Unreliable integrations are surfaced as an owner's problem, which is the only thing that ever gets them fixed.
- Case state means what it says, which is what makes the subject's confirmation truthful.

**What it costs**

- Blocked cases accumulate for targets with weak integrations, and the backlog is uncomfortable to look at.
- There is standing pressure to reclassify silence as pending, and it will come with a business case each time.
- Escalation requires every target to have a current named owner, which is an organisational dependency the platform cannot enforce.

**Choose differently when.** If contracts with recipients mandated machine-readable attestation with penalties, the silent case becomes rare enough that its handling matters less. Nothing about the technical design would change; the volume would.

**Why it holds up over time.** Treating absence of response as failure rather than success is a first principle of any protocol with an untrusted counterparty, and does not depend on the domain.

> **Lesson.** Decide early which direction an unknown resolves in, and pick the direction whose failures are visible. A system that resolves unknowns optimistically is a system that lies to its owners.

#### ADR-13 · The suppression list is mandatory on every ingestion and restore path

**Status:** Accepted  ·  **Shown on views:** 11, 14, 22

*What stops an erased subject from coming back through a backup restore, a replayed event stream or a vendor re-upload?*

**Context.** Erasure is a point-in-time act, and data flows continuously. A restore from a four-week-old backup reintroduces a subject erased three weeks ago. A replayed stream re-creates rows. A vendor's periodic file upload re-adds a contact the vendor was told to delete and did, from a copy they kept for reconciliation. Each of these is a separate re-entry path and there is no single system that owns all of them — which is precisely why a per-path solution fails: there is always one more path.

**Decision.** A single compact suppression list of hashed erased subject keys and erasure dates is maintained per region, and consultation is a mandatory precondition of every ingestion and restore path in the estate. A path that cannot consult it cannot be registered as a processing target. The list is retained indefinitely, because the obligation does not expire.

**How it is realised on AWS.** A DynamoDB table of hashed subject keys with erasure dates, replicated in-region, with a compact snapshot published to S3 and CloudFront for paths that need a local copy. The SDK exposes a suppression check alongside the decision call; the warehouse applies it as a filter on load; restore runbooks require a suppression pass before a restored dataset is made readable. Non-consultation is detected by the verification prober, which probes restored and re-ingested data specifically.

| Option | Verdict | Reasoning |
|---|---|---|
| One mandatory suppression list, consulted on every entry path | Chosen | One control that holds across all re-entry paths, including the ones nobody enumerated; costs a check in the hottest write paths. |
| Re-run erasure after each restore | Rejected | Correct in principle and depends on someone remembering, which is the failure mode being designed against. |
| Per-subject encryption so backups self-shred | Rejected | Elegant where it applies, and it does not cover a vendor re-upload or a replayed stream carrying plaintext. |
| Shorter backup retention to bound the window | Right elsewhere | Genuinely effective and usually impossible: backup retention is set by recovery and legal requirements, not by privacy. |

**What it buys**

- One control covers restore, replay and re-upload, including paths discovered after the design was written.
- The list holds only hashed keys and dates, so it is cheap to replicate and carries minimal content risk.
- Non-consultation is independently detectable, rather than relying on each path's owner to confirm it.

**What it costs**

- A check on every ingestion path is a hot-path cost imposed on systems that get no benefit from it.
- The list is pseudonymous personal data retained indefinitely, which needs its own lawful basis and its own explanation to subjects.
- It grows forever, and the hash is only as protective as the salt management around it.

**Choose differently when.** If the estate's ingestion consolidated behind a single platform-owned pipeline, the check moves to one place and stops being a distributed obligation. If a jurisdiction held that retaining erased subjects' hashed keys is itself unlawful, the design would need per-subject crypto-shredding to carry the whole burden, which it cannot for plaintext re-uploads.

**Why it holds up over time.** Any system that erases from a continuously-fed store needs a negative index. The shape of that requirement does not change with technology.

> **Lesson.** You cannot enumerate every way data gets back in. Put the control where everything enters rather than where you believe things enter.

### Residency and evidence

*Where the record of permission lives, and what has to survive the request to be forgotten.*

#### ADR-14 · Residency is enforced in the data path, and consent metadata is personal data

**Status:** Accepted  ·  **Shown on views:** 08, 16, 20

*Is a jurisdictional boundary a property of the architecture or a column in the database?*

**Context.** One global deployment with a residency tag per row is dramatically cheaper to run and is a single configuration error away from an unlawful transfer — and that error is undetectable from inside the system, because the data served looks correct. Isolated regional stacks make the boundary structural: there is no replica to read from, so a routing mistake fails rather than leaks. The cost is nine deployments, nine key hierarchies, no cross-border failover, and the acceptance that a region's bad day denies its own subjects' consent-based purposes. The subtler point is that the consent record is itself personal data about the subject, so the boundary applies to this platform's own stores and not only to the data it governs.

**Decision.** Each jurisdiction is served by its own regional stack holding the ledger, projection, identity index, case state, suppression list, keys and audit for its subjects. Nothing crosses: no replica, no key, no failover. Only the purpose registry and the signed policy bundle — which hold no personal data — replicate globally. A request carrying a subject key outside the region's jurisdiction is rejected or referred, never served from a replica that happens to hold it.

**How it is realised on AWS.** Regional stacks in eu-west-1/eu-central-1, us-east-1/us-west-2 and ap-south-1, each with its own DynamoDB tables, KMS CMK with no cross-region grant, Step Functions, and S3 Object Lock audit bucket. Aurora Global carries the registry. A jurisdiction router at the edge, backed by Route 53 and edge policy, directs each subject's traffic to their boundary, and the regional API rejects foreign subject keys outright rather than proxying them.

| Option | Verdict | Reasoning |
|---|---|---|
| Isolated regional stacks; only definitions global | Chosen | Makes the boundary structural so a routing bug fails instead of leaking; costs nine stacks and no failover. |
| One global deployment with residency tags and policy routing | Rejected | Far cheaper, and a single misconfiguration produces an undetectable unlawful transfer. |
| Global control plane with regional data planes, personal data regional | Rejected | Essentially the chosen design with a larger global surface; rejected because any global component holding subject keys reintroduces the risk. |
| Sovereign operator-run deployment per jurisdiction | Right elsewhere | Right where a jurisdiction requires operational sovereignty too; here it multiplies cost without changing the data boundary. |

**What it buys**

- An unlawful transfer requires a deliberate act, not a configuration mistake.
- Adding a jurisdiction is a new deployment and new configuration, not a schema change.
- Per-region keys mean a compromise of one boundary's key hierarchy cannot decrypt another's.

**What it costs**

- A region's unavailability denies its subjects' consent-based purposes, and there is no failover to offer.
- Nine stacks multiply operational cost, patching surface and configuration drift risk.
- The jurisdiction router becomes the one component whose failure is an unlawful transfer, so it is now the most safety-critical routing decision in the estate.

**Choose differently when.** If a jurisdiction permitted consent metadata to be held outside its boundary under a recognised transfer mechanism, the stacks could consolidate and the cost falls sharply. More likely the pressure runs the other way: new jurisdictions are added, which this design absorbs and the global-with-tags design does not.

**Why it holds up over time.** Jurisdictional fragmentation of data protection is a decade-long trend. A design whose boundary is enforced by routing and keys scales with that trend; one whose boundary is a policy document does not survive its first audit.

> **Lesson.** If a boundary matters legally, make it impossible to cross rather than wrong to cross. A boundary enforced by correctness of configuration is a boundary you will cross by accident.

#### ADR-15 · Evidence that an erasure happened survives the erasure

**Status:** Accepted  ·  **Shown on views:** 06, 12, 22

*When a subject asks to be forgotten, may the platform keep the record proving it forgot them?*

**Context.** The obligation to erase and the obligation to demonstrate compliance point in opposite directions. Honouring erasure completely would destroy the evidence that erasure occurred, leaving the organisation unable to answer the regulator's most basic question and unable to defend itself against a claim that it did nothing. Keeping full records means holding personal data about someone who explicitly asked not to be held. There is no technical resolution: the tension is real, and the only question is which way it is resolved and whether the subject is told.

**Decision.** The platform retains the minimum evidence needed to prove an erasure occurred — subject key, case identifier, per-target outcomes, timestamps, the technique applied — under a lawful basis of legal obligation that the subject cannot withdraw. Everything beyond that minimum is erased with the rest. The retention and its basis are stated plainly to the subject in the erasure confirmation, not buried in a policy.

**How it is realised on AWS.** Audit records land in S3 with Object Lock in compliance mode, hash-chained so a gap is detectable, with independent credentials and a separate blast radius from the operational stores. The erasure case removes the subject's consent entries' content while retaining the case skeleton and outcomes. The retained minimum is enumerated in the records of processing and reproduced verbatim in the subject's confirmation.

| Option | Verdict | Reasoning |
|---|---|---|
| Minimum evidence retained under legal obligation, stated to the subject | Chosen | Defensible and honest; it still means holding data about someone who asked not to be held. |
| Erase everything including the evidence | Rejected | Maximally faithful to the request and leaves the organisation unable to prove it complied, including against a false claim that it did not. |
| Retain full case detail indefinitely | Rejected | Convenient for audit and keeps far more than proving compliance requires, which is the definition of excessive processing. |
| Hand the evidence to the subject and keep nothing | Right elsewhere | Appealing and unworkable: the organisation's obligation to demonstrate compliance cannot be discharged by data only the subject holds. |

**What it buys**

- The regulator's question is answerable for erased subjects, which is otherwise the one population about whom nothing can be said.
- The retained set is enumerable and minimal, so the trade is reviewable rather than open-ended.
- Hash-chained write-once storage means a missing record is detectable, which is what makes the evidence worth anything.

**What it costs**

- The platform holds personal data about people who asked not to be held, under a basis they cannot withdraw. That is the honest cost and it cannot be designed away.
- The subject's confirmation must explain this, which makes a good-news message partly bad news.
- Seven-year write-once retention of audit records is a cost and a disclosure obligation of its own.

**Choose differently when.** If a regulator specified a shorter or narrower evidentiary minimum, the retained set shrinks accordingly — the design already enumerates it, so this is configuration. If demonstrating compliance were accepted on aggregate rather than per-subject evidence, the retained set could become statistical and the tension largely disappears.

**Why it holds up over time.** The conflict between a right to erasure and a duty to demonstrate compliance is structural to every accountability regime, and will be present in whatever replaces the current ones.

> **Lesson.** When two obligations genuinely conflict, resolve it explicitly, minimise what you keep, and tell the person. A conflict resolved silently is one you will be asked about under worse conditions.

## Every package used, in one table

Terms used precisely in this package, including several that are used loosely elsewhere. Where a looser alternative exists, it is named along with what it costs.

| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Purpose | A declared, versioned reason for processing personal data, carrying data categories, retention, permitted systems and recipients. | The unit everything else is keyed on: consent, decisions, erasure fan-out and cost all hang off it. | "Use case" or "processing activity", which blur the versioned legal object with the product feature that happens to rely on it. |
| Lawful basis | The legal justification for processing under a purpose, declared per purpose version per jurisdiction. | Decides the default: no processing before a grant under consent, processing until objection under legitimate interest. | Treating consent as the universal basis, which forfeits lawful processing and trains users to click through everything. |
| Widening change | A purpose change adding a data category, a recipient, or a longer retention period. | Triggers a new purpose version and fresh permission where the basis is consent; a narrowing change applies immediately. | "Update", which lets a grant given for something small silently cover something larger. |
| Propagation ceiling | The published maximum time between a withdrawal being committed and every enforcement point reflecting it — 15 minutes here. | The real compliance bound. Within it, staleness is accepted; beyond it, the platform is in an incident. | A latency target, which implies a performance concern rather than an accruing legal exposure. |
| Staleness ceiling | The maximum age a local cached consent entry may reach before the purpose's failure posture applies. | Converts a decision-plane outage into a bounded, declared degradation instead of silent divergence. | A TTL, which is the same mechanism without a stated obligation attached to it. |
| Failure posture | A per-purpose registry field deciding whether an unanswerable decision fails closed or open. | Makes the dangerous choice explicit, approved by two people, visible in the records of processing, and audited whenever relied upon. | A global fallback, which either makes this platform the availability ceiling for the estate or legalises an outage's worth of processing. |
| Crypto-shredding | Satisfying erasure by destroying the subject's encryption key, leaving unreadable ciphertext. | The only erasure technique available for append-only and integrity-protected stores. | Calling it deletion, which overstates what happened and may not satisfy every jurisdiction. |
| Attested versus verified | Attested is the target's claim that it deleted; verified is the platform's own evidence that the data is gone. | Keeping them distinct is what stops a case closing on a claim by the party with the least incentive to check. | "Confirmed", which collapses the two and makes every completeness statement unfalsifiable. |
| Suppression list | The set of hashed erased subject keys consulted by every ingestion and restore path. | The single control preventing resurrection by backup restore, stream replay or vendor re-upload. | Re-running erasure after each restore, which depends on somebody remembering every path. |
| Jurisdictional boundary | A set of regions serving one jurisdiction's subjects, with no replica, key or failover crossing it. | Makes residency a structural property rather than a configuration value that can be set wrongly. | A residency column, which is a single misconfiguration away from an undetectable unlawful transfer. |
| Subject key | The pseudonymous internal identifier for a data subject, to which consent, cases and evidence are keyed. | Lets the platform govern a person without holding their identifiers or their data. | Using an email address or user id, which spreads a directly identifying value through every store here. |
| UNKNOWN | A decision verdict meaning the subject's permission state could not be determined. | Separates the system working (DENY, no grant) from the system broken, which otherwise look identical in the metrics. | Folding it into DENY, which makes a failing projection indistinguishable from a population that all opted out. |
