An interviewer says "we are going federated across about 40 domains. Which governance policies do we put in the platform as code and which stay as written guidance? Talk me through how you decide, and what the code version costs us."
Show the full answer Hide the answer
What the interviewer is testing
Whether you can separate a policy that is decidable from metadata at a specific moment from one that requires judgement, and whether you understand that encoding a policy moves it from the governance budget to the platform's on-call budget. The weak candidate lists the four data mesh principles. The strong one produces a test that sorts any proposed policy into one of two piles in front of the interviewer.
The clarifying questions that change the answer
How many domains, and do they share one platform or several? A policy as code assumes a single enforcement point; at two platforms you are writing it twice and they will diverge. Is there an external auditor who needs evidence, which raises the value of machine-generated proof? And what is the current violation rate, because a policy encoded before anyone has measured how often it is broken is a blocked build with no baseline.
A strong answer's arc
A policy goes in the platform when three things hold: it is decidable from metadata available at build or query time, it has a remedy the producing team can carry out themselves, and its false-positive rate is low enough that blocking is proportionate. Examples that pass all three: every published dataset resolves to an owner group in the directory; every column labelled above Internal has a masking or access policy attached; a column deletion fails the producer's build unless that column has been marked deprecated for at least 30 days.
What stays as guidance is what needs a person with authority: what a metric means, whether a particular join is an appropriate use of the data, whether a seven-year retention is proportionate. Guidance without a bounded escalation path is not a policy, it is a wish. Each one needs a named decider and a service level, for example five business days, or domains will simply proceed.
What a strong answer adds
The cost, stated plainly. Every computational policy becomes a build-time dependency on the platform, so the platform team now owns an availability problem for 40 domains' pipelines. At a 2% false-positive rate across 40 domains deploying daily, that is close to one wrongly blocked build every day, and the organisational response to a daily wrong block is an exception mechanism that becomes the normal path within a quarter.
So ship every new policy in warn mode first, publish the violation count by domain, and promote it to blocking only when the rate falls below a stated threshold such as one in 200 builds. Give every exception an owner and an expiry date, and report the count of open exceptions to the same forum that approved the policy. A policy with 60 permanent exceptions is not enforced; it is theatre with a dashboard.
Common weak answers
- "Everything as code." This encodes semantics that cannot be decided from metadata, so the platform starts failing builds over definitional disagreements it cannot resolve.
- "A central board approves each dataset." That is the bottleneck federation existed to remove, reintroduced at the point of publication.
- "The domains decide their own policies." Then there are 40 incompatible sensitivity schemes and no cross-domain join is safe, which is the failure mode federated governance is named for.
When this is the wrong answer
With three or four domains and one platform team, encoding policy is slower than a conversation. Write the rules down, review at publication, and revisit when the number of domains passes the point where the platform team no longer knows the producers by name.