intermediate 2 min answer

A group holds customer data in separate EU and US entities, and the rows may not be copied across the border. The board wants one revenue dashboard. The team puts a federation layer over both regions and queries in place rather than consolidating. What has that bought, when does the bill arrive, and how do they keep the option to reverse it?

federationresidencypushdowncross-regionreversibility
Show the full answer Hide the answer

What is gained

The copy that would have needed a legal basis never exists. No transfer assessment, no second retention clock, no duplicate deletion obligation when a customer exercises erasure, and one fewer store to secure and audit. The dashboard is also live rather than a day behind, which is a genuine product improvement and is the reason the team will defend the choice after it starts hurting.

What is paid

  • Latency floors at the slowest link. A transatlantic round trip is roughly 70-90 ms one way. A plan that makes several sequential round trips per source turns a 200 ms query into seconds, and no amount of local caching changes the floor.
  • Joins do not push down across sources. The engine must ship one side. A 50,000-row dimension is fine; a 40-million-row fact is not. The only shape that works is aggregate inside each region and combine the results - which constrains what the dashboard can offer, permanently.
  • The sources carry analytical load they were not sized for. This is the failure that matters: a federated scan against an operational replica raises replica lag, and lag reaches the application. The bill arrives as an operational incident, not a slow report.
  • No result history. When a number changes between Monday and Tuesday, there is no stored version to compare, so every question becomes an investigation.

When the cost becomes visible

At three predictable moments: when someone asks for row-level drill-through behind an aggregate; when concurrency passes a few dozen users, because each query holds connections in every source for its whole duration; and at quarter close, when both regions are busiest and the operational databases have least headroom. Each is a growth event, not a defect, which is why the layer looks fine in the pilot.

How to keep the option to reverse

Define the dashboard as a contract over per-region aggregates. Each region publishes a small table - revenue by day, product and segment - containing no personal data and therefore free to move. Today the federation layer reads those two tables and unions them. If latency or load becomes unacceptable, the same aggregates are pushed to a central store on a schedule, and the dashboard does not change. That is a scheduling change instead of a redesign, and it costs nothing to set up now.

Also: give each federated source an explicit load budget - a concurrency cap, a row cap and a statement timeout - and point the connector at a replica dedicated to analytics. Without it, the first unbounded query is an operational incident.

When not to federate

When the aggregates can legally move, materialise them. Federation is for data that cannot move and queries that are few. If the only reason for the federation layer is that building two small export jobs felt like duplication, the copy is the simpler system and it is faster, cheaper per query and explainable after the fact.