Data Never Leaving the Phone Doesn't Mean the Privacy Problem Is Solved
Having large numbers of phones jointly train one model, with raw data staying on each device and only the learned updates uploaded — this kind of approach is called federated learning.
ImportanceMaterialEvidenceE3 inspectableWrite-upQuick
Having large numbers of phones jointly train one model, with raw data staying on each device and only the learned updates uploaded — this kind of approach is called federated learning. A long 2019 survey listed the problems this approach has yet to solve: unreliable devices, differing data across participants, constrained bandwidth, hard-to-prove privacy — and these are all entangled with one another.
Today, AI offerings in input methods, healthcare, and finance love to pitch with one line: data never leaves the device, so privacy is solved. Later research is still working through the items this survey listed one by one, and no single method solves them all. When you hear that pitch, press with the checklist: what about differing data across participants? How do you prove nothing was actually peeked at?
If you want to compare the merits or performance numbers of specific methods, don't use it to judge — it ran no experiments, only an inventory of problems. And if your scenario is learning without a central server, where devices connect directly to each other, the problems it lists don't match either.
Advances and Open Problems in Federated Learning (2019) | Next review 2027-09-20