It's My Data Too: Private ML for Datasets with Multi-User Training Examples
Arun Ganesh, Ryan McKenna, Hugh Brendan McMahan, Adam Smith, Fan Wu
摘要
We initiate a study of algorithms for model training with user-level differential privacy (DP), where each example may be attributed to multiple users, which we call the multi-attribution model. We first provide a carefully chosen definition of user-level DP under the multi-attribution model. Training in the multi-attribution model is facilitated by solving the contribution bounding problem, i.e. the problem of selecting a subset of the dataset for which each user is associated with a limited number of examples. We propose a greedy baseline algorithm for the contribution bounding problem. We then empirically study this algorithm for a synthetic logistic regression task and a transformer training task, including studying variants of this baseline algorithm that optimize the subset chosen using different techniques and criteria. We find that the baseline algorithm remains competitive with its variants in most settings, and build a better understanding of the practical importance of a bias-variance tradeoff inherent in solutions to the contribution bounding problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Practical and Private (Deep) Learning Without Sampling or ShufflingPeter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar 等ICML 2021 · 被引用 239 次
- Improved Differential Privacy for SGD via Optimal Private Linear Operators on Adaptive StreamsSergey Denisov, H. Brendan McMahan, John Rush, Adam D. Smith 等NeurIPS 2022 · 被引用 96 次
- Multi-Epoch Matrix Factorization Mechanisms for Private Machine LearningChristopher A. Choquette-Choo, Hugh Brendan McMahan, J. Keith Rush, Abhradeep Guha ThakurtaICML 2023 · 被引用 62 次
- Shifted Inverse: A General Mechanism for Monotonic Functions under User Differential PrivacyJuanru Fang, Wei Dong, Ke YiCCS 2022 · 被引用 13 次
- Time-Aware Projections: Truly Node-Private Graph Statistics under Continual ObservationPalak Jain, Adam Smith, Connor WagamanS&P 2024 · 被引用 11 次
相关 Paper
- Smoothly Bounding User Contributions in Differential PrivacyAlessandro Epasto, Mohammad Mahdian, Jieming Mao, Vahab S. Mirrokni 等NeurIPS 2020 · 被引用 17 次
- Algorithms for bounding contribution for histogram estimation under user-level privacyYuhan Liu, Ananda Theertha Suresh, Wennan Zhu, Peter Kairouz 等ICML 2023 · 被引用 14 次
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale 等NeurIPS 2021 · 被引用 113 次
- User-Level Differential Privacy With Few Examples Per UserBadih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi 等NeurIPS 2023 · 被引用 19 次
- Task-aware Privacy Preservation for Multi-dimensional DataJiangnan Cheng, Ao Tang, Sandeep ChinchaliICML 2022 · 被引用 8 次
