advanced 3 min answer

A product lead at a super-app of Tencent/WeChat's shape says the new on-device keyboard model can ship without a consent flow because federated learning means the data never leaves the phone. You have ten minutes at the architecture review board. What do you say and what would you actually build?

federated-learningsecure-aggregationdifferential-privacyon-device-mllawful-basis
Show the full answer Hide the answer

What the board is really asking

Whether you understand that a privacy-enhancing technology changes the risk profile of a design and not the question of whether you have a basis to process. And whether you can price the engineering, because federated learning is presented as free and is not.

The clarifying questions that change the answer

What exactly is the model trained on? A keyboard sees passwords typed into the wrong field, card numbers and one-time codes. How many clients per training round? A cohort of 50 is not a cohort. Is any aggregate published externally, or is the model internal only? And what is the fallback for devices that cannot participate, since that fallback is usually a server-side path that does collect raw data?

The correction

A model update computed from someone's keystrokes is derived from their personal data, and stays personal data unless something makes it non-attributable. "Deep Leakage from Gradients" (Zhu, Liu and Han, NeurIPS 2019) reconstructs original training samples from shared gradients in a small number of optimisation steps, on both vision and language tasks. A design in which the server sees one client's raw update has not moved the data off the phone in any sense that matters; it has changed its representation.

What you would actually build

  1. Secure aggregation, so the server sees only the sum across a cohort and never one client's update. The practical protocol is Bonawitz et al., ACM CCS 2017, and the property to check is dropout resilience, because mobile clients vanish mid-round as a matter of course.
  2. Client-level differential privacy: clip each update to a norm bound, add noise calibrated to the cohort size, and run a privacy accountant across rounds. The guarantee holds over the whole training run, so the budget is spent once and cannot be recovered by rephrasing the question later.
  3. An on-device contribution filter that never learns from secure text entry, numeric-only fields or anything pasted from a password manager. This is the cheapest control of the three and the one most often left out.

What it costs

Cohorts in the thousands per round for the noise to be tolerable. Wall-clock training in days to weeks rather than hours, because rounds wait for devices that are charging, idle and on wifi. No ability to inspect a bad training example, so debugging is done against synthetic proxies and a regression can take a release cycle to find. A second serving path for devices that cannot participate. Budget a small team for two quarters before the first model ships.

Common weak answers

"It is anonymous because it is aggregated" — aggregation over a cohort of one is not aggregation, and cohort size is the parameter nobody checks. "We add noise" without an accountant, which spends an unbounded budget. And treating secure aggregation as a substitute for differential privacy: it hides an individual update from the server, not from the trained model, which can still memorise a rare string contributed by one person.

When this is the wrong answer

If consent can be obtained and the data centralised, centralised training is an order of magnitude cheaper, faster and debuggable, and the honest recommendation is to build the consent flow. Federated learning earns its cost only when centralising is genuinely unavailable: a signal that exists only on the device, a consortium of hospitals that cannot pool records, or a transfer a regulator forbids. Choosing it to avoid a consent dialogue buys the cost without the justification.