Two hospital groups want to measure whether a treatment pathway improves outcomes, using patient records neither may disclose to the other. The analysis needs per-patient linkage across both datasets. Which approach fits?
Show the full answer Hide the answer
The deciding property
Whether raw records must move, and who must never see them. The constraint is that neither party may disclose records to the other, and the analysis needs per-patient linkage. Those two requirements together eliminate every option that creates a pooled dataset in anybody's hands.
Federated computation satisfies both: the computation travels to each dataset, each party computes over its own records, and only aggregates leave. Secure aggregation means the coordinator sees the combined result without seeing either party's contribution.
Why federated computation with secure aggregation
- No raw record crosses a boundary, which is the requirement as stated rather than a mitigation of it.
- Linkage can be done on a privacy-preserving identifier — a salted, jointly derived pseudonym — so matched patients are counted without either side learning which of its patients are in the other's cohort beyond the linkage itself.
- Secure aggregation removes the coordinator from the trust boundary, which matters because the coordinator is often a third party such as a research network.
- Add a differential-privacy budget on top of the published aggregates, because with small cohorts an aggregate can identify an individual. The two techniques are complementary: federation controls where data sits, differential privacy controls what the outputs reveal.
The costs are substantial and should be stated: a shared protocol and schema both sides implement, agreement on the exact analysis before it runs (ad hoc exploration is nearly impossible), a joint governance process, and accuracy loss from the privacy budget. Expect months of legal and protocol work against days of analysis.
Why the others fail
- Private set intersection then exchange the matched records does the first half correctly and then breaks the constraint. PSI is the right tool for learning the size or membership of an overlap without disclosing non-members — two retailers computing a shared-customer count — but exchanging the matched records is precisely the disclosure that is forbidden, and the fact that they are matched makes it more sensitive rather than less.
- Differential privacy on a pooled dataset solves the wrong problem. It protects the outputs of an analysis, and it requires someone to hold the pooled raw data first. That holder is a new concentration of patient records and a new legal basis nobody has.
- Synthetic data from each side is appropriate for building and testing a pipeline, and it cannot support this conclusion. Synthetic generation preserves the statistical properties the generator was told to preserve and destroys the per-patient linkage entirely, so a joint cohort analysis cannot be performed. Treating synthetic data as a substitute for real analysis is the most common misuse of it, and the result is a finding about the generator.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| Only the overlap size is needed | Private set intersection | Far simpler and sufficient for the question |
| A lawful joint controller or trusted third party exists | Pooled analysis in a secure enclave with a DP budget | Removes protocol complexity; the trust question is answered legally |
| The question is a population statistic with no linkage | Independent analyses plus meta-analysis | No cross-silo machinery needed at all |
| You are building the pipeline, not answering the question | Synthetic data | Correct tool for development and testing |
When not to reach for any of this
If a lawful basis and a willing trusted third party exist, a secure enclave with a pooled dataset and a published privacy budget is simpler, faster and easier to explain to an ethics committee than a federated protocol. Privacy-enhancing technologies substitute cryptography for trust, and they are worth their cost only where the trust genuinely cannot be established. Where it can, the simpler architecture is also the more auditable one.