advanced 3 min answer

An e-commerce group of Flipkart's shape runs an in-India cell for regulated payment data and an EU cell for everything else. A quarterly control test samples 500 rows from the EU warehouse and finds 38 carrying Indian cardholder identifiers. The routing layer is correct and the replication topology is correct. Where do you look and in what order?

data-residencytelemetry-leakagedata-classificationegress-controldiagnosis
Show the full answer Hide the answer

Read the ratio first

38 of 500 is 7.6%. If India is 40% of order volume and the main write path were misrouting, the contaminated share would sit near 40%, not near 8%. A contamination rate well below the traffic share says the leak is on a minority path, and minority paths in an order system are error paths, admin paths and batch paths. That single arithmetic step removes the routing layer and the replication topology from the list before anyone reads a config file, which is why the ratio is the first thing to compute.

The first three places to look, in order

  1. Error and crash telemetry. Exception payloads, breadcrumbs and request-body captures are serialised by a library that knows nothing about the routing rules, and they go wherever the observability backend lives. An error rate of a few per cent of orders matches the ratio almost exactly. Check the collector's field allow-list and whether it exists at all.
  2. Back-office and support tooling. A unified search that queries both cells and caches results, or an export button that writes a CSV to shared object storage in the EU. Support tools are built after the split, by a different team, against a union view.
  3. Batch extracts and BI jobs written before the split, still pointing at a view that unions the two stores. These run nightly and nobody has looked at their definition in two years.

Then check derived keys: an idempotency key or a dedup hash computed from a card identifier is global by design, and it carries the identifier's information into every cell that uses it.

The misleading signal

The routing layer's own dashboards are green and correct, and they will hold the investigation for a day if you let them. They are green because they measure the path that was designed, and every path that failed here is one nobody classified as a data path at all. A control that only watches the designed path cannot detect a leak on an undesigned one.

The fix, and the alert that would have caught it

Classification at the field level in a schema registry, then an egress check at the EU ingest boundary that rejects any row whose declared classification is in-India-only, applied to the telemetry collector on the same terms as the application. Choose a default-deny check over a pattern blocklist unless the field set is tiny and stable, since a blocklist silently misses the field a new release adds. Replace the quarterly 500-row sample with a continuous scanner over the EU store matching the regulated-field patterns, alerting on the first hit. The sample was never adequate: 500 rows detects contamination around 0.6% or higher with any confidence, and a leak an order of magnitude smaller than this one would have passed four tests in a row.

When this is the wrong amount of work

Get the obligation's wording first. Some storage rules permit a copy abroad for the foreign leg of a cross-border transaction, and some cover only a defined set of fields rather than the whole record. If the rule is field-scoped, the correct response is a field classification and one egress check, not a programme to close every path. Spending a quarter hardening paths that the obligation allows is a real cost with no compliance return, and it is the more common mistake of the two.