advanced 3 min answer

A consumer platform with 40 million accounts receives about 600 subject access requests a month. Honouring one means assembling a person's data from 23 services that each own their own store. Estimate what the manual path costs and say what changes the number by an order of magnitude.

dsargdprsubject-indexdata-lineageestimation
Show the full answer Hide the answer

The assumptions, stated

  • 600 requests a month, roughly 30 per working day.
  • 23 owning services. Three hold most of a person's data and take about two hours each to locate, extract and review. Twenty hold a little and answer from a runbook in about fifteen minutes.
  • Review is not optional: an export that includes another person's data is a second breach, so somebody reads it.

The arithmetic

Per request: 3 × 2 h + 20 × 0.25 h = 6 + 5 = 11 person-hours.

Per month: 600 × 11 = 6,600 person-hours. At roughly 160 productive hours a month that is about 41 full-time engineers, spread as a tax across 23 teams rather than appearing in anyone's headcount.

Range: between about 20 and 80 FTE depending on the per-service minutes.

What the number rules out

The manual path, immediately and unambiguously. No organisation staffs 41 engineers on access requests, so what actually happens is that requests queue, the statutory deadline is missed, and teams start answering from whatever store is easiest to query. That last part is the real failure: the analytics warehouse holds a copy, not the authoritative set, so an answer built from it is incomplete and demonstrably so the moment a regulator asks what else you hold.

It also rules out the opposite extreme. A single service that owns all personal data would make this trivial and is not a thing you can retrofit onto 23 teams.

What changes it by two orders of magnitude

A subject index: a registry written at the same time as the data, mapping a canonical subject identifier to tuples of service, dataset and record locator. Assembly becomes a fan-out of automated reads against locators somebody already recorded.

Build cost: every producing service publishes locator rows on write, roughly an engineer-month each, so about two engineer-years across 23 services. The index itself is small. At 40 million subjects and a handful of locators each, a few hundred million rows of about 100 bytes is tens of gigabytes, which is not a design constraint. Against 41 FTE of recurring cost, payback is immediate.

The same index is what makes deletion verifiable rather than hopeful, which is the larger prize and the reason to build it before anyone asks.

Which assumption dominates the error

The per-service minutes, which vary by a factor of ten depending on whether a service can query by subject at all. The services where the number explodes are the ones with no subject-keyed index, typically event logs and free-text stores, and they are also the ones a locator registry helps most. Measure three real requests end to end before quoting any figure.

When this is the wrong answer

Below roughly a hundred requests a month across a handful of services, a runbook with named owners and a tracked deadline is correct, and the index is premature. The threshold to watch is not the request count on its own but whether measured p95 assembly time is approaching the statutory response deadline, which is one month under GDPR and extendable to three. When p95 crosses about half the deadline, start building, because the queue is what fails first.