protocol

Secure Aggregation

also called Private Aggregation, Cohort-Level Aggregation

A cryptographic protocol that lets a server learn only the sum of many clients' model updates and never any single client's update, which is what makes an on-device training claim survive scrutiny.

federated-learningcryptographydifferential-privacyon-device-mlprivacy

A team ships on-device training and tells the review board that no personal data leaves the phone. The server still receives one gradient vector per device per round. "Deep Leakage from Gradients" (Zhu, Liu and Han, NeurIPS 2019) reconstructs original training samples from shared gradients in a small number of optimisation steps, on both vision and language tasks. The raw data did not leave, and a close enough copy of it did.

Secure aggregation closes that gap. Clients mask their updates with pairwise secrets that cancel when summed, so the server can compute the cohort total and cannot read any individual contribution. The output is arithmetically identical to plain summation; the difference is what the server learns on the way.

The protocol that made it practical is Bonawitz et al., ACM CCS 2017. Its contribution is not the masking idea but the engineering around it: clients drop out constantly on mobile networks, and a naive pairwise masking scheme breaks when any participant disappears mid-round.

Why it matters

Without it, a federated design is a distributed collection pipeline with extra steps. The lawful basis question is unchanged, because a model update derived from someone's keystrokes is still derived from their personal data, and the breach exposure is worse than a central database in one respect: the updates arrive continuously and are rarely retained under the same controls as the warehouse.

With it, the smallest unit the server ever observes is a cohort, and cohort size becomes the parameter the whole privacy argument rests on. That is a property you can state to a regulator, test in code, and alert on.

Implementation patterns

  • Pairwise masking with secret sharing. Each pair of clients derives a shared mask that one adds and the other subtracts. Masks cancel in the sum. Each client also shares its own secrets with a threshold of peers so a dropped client's mask can be reconstructed and removed.
  • Dropout resilience as a first-class requirement. Set the reconstruction threshold from the observed dropout rate, not from an ideal. On consumer mobile, a third of selected clients failing to complete a round is unremarkable.
  • A minimum cohort size enforced server-side. Refuse to aggregate below it. The common figure in production designs is in the low thousands per round, because the noise added later is calibrated to cohort size.
  • Differential privacy on top, not instead. Clip each update to a norm bound, add noise, and track the budget with an accountant across rounds. Secure aggregation hides an individual update from the server; it does nothing about what the trained model memorises.
  • An on-device contribution filter that never learns from secure text entry, numeric-only fields, or clipboard pastes. This is the cheapest of the three controls and the one most often skipped.

Industry example

Mobile keyboard prediction is the canonical published use: rounds run on devices that are charging, idle and on wifi, with a server that aggregates thousands of masked updates per round. The pattern recurs in hospital consortia, where no participant may send records to a central store but all want one model, and in browser telemetry where the vendor wants counts without per-user histories.

Failure scenarios

  • A cohort of one. Client selection during a rollout in a small market produces a round with three participants, and the sum is effectively an individual update. Nothing errors.
  • Dropout below the reconstruction threshold, so the round cannot be unmasked and is discarded. Training stalls and the symptom looks like a model quality problem.
  • Secure aggregation treated as a substitute for differential privacy, after which the trained model reproduces a rare string that exactly one user typed.
  • A parallel debug path that logs an unmasked update for one device to diagnose a bug, and stays on.

Trade-offs

Choose it Gains Pays
Secure aggregation The server never sees an individual update, so the privacy claim is structural Cohorts in the thousands · rounds that wait for device availability · no ability to inspect a single bad example
Plain federated averaging Simple and debuggable The server sees every client's gradient and the design leaks
Centralised training An order of magnitude cheaper and fully debuggable Needs a lawful basis for the transfer and a real central store to protect

Wall-clock training moves from hours to days or weeks, because rounds wait for devices rather than for GPUs. Debugging costs the most in practice: when a model regresses there is no example to look at, so teams build synthetic proxies and a regression can take a release cycle to locate.

When not to use it

If the data can be centralised under a lawful basis the organisation already has, centralise it. Centralised training is cheaper, faster and debuggable, and choosing federated learning to avoid writing a consent dialogue buys the whole cost without the justification. Secure aggregation also adds nothing where the aggregate itself is the disclosure risk: publishing a count of 4 with noise is a differential privacy problem, and hiding who contributed the 4 does not fix it. The rule: secure aggregation protects the input, differential privacy protects the output, and you need whichever one matches the thing you are exposing.

Interview question

Q: Your team proposes federated learning with secure aggregation for a keyboard model. Tell me what you would set the minimum cohort size to and how you would defend that number, then tell me what secure aggregation still does not protect you from.

What a strong answer covers: cohort size derived from the differential privacy noise calibration rather than picked; a server-side refusal below it; dropout rate observed and the reconstruction threshold set from it; and the clear statement that memorisation by the trained model is untouched by aggregation and needs clipping, noise and an accountant.

Quick check

Quiz: A federated round completes with 12 participating clients. Why is that a privacy incident even though secure aggregation worked correctly? Because the sum over 12 clients approximates an individual contribution, and the protocol guarantees only that the server cannot see one update, not that the aggregate conceals one.

Flashcard: Secure aggregation hides X from the server but not Y from the model. X is any individual client update; Y is a rare training string the model memorises, which needs client-level differential privacy with a budget tracked across rounds.