It's My Data Too: Private ML for Datasets with Multi-User Training Examples
Arun Ganesh, Ryan McKenna, Hugh Brendan McMahan, Adam Smith, Fan Wu
Abstract
We initiate a study of algorithms for model training with user-level differential privacy (DP), where each example may be attributed to multiple users, which we call the multi-attribution model. We first provide a carefully chosen definition of user-level DP under the multi-attribution model. Training in the multi-attribution model is facilitated by solving the contribution bounding problem, i.e. the problem of selecting a subset of the dataset for which each user is associated with a limited number of examples. We propose a greedy baseline algorithm for the contribution bounding problem. We then empirically study this algorithm for a synthetic logistic regression task and a transformer training task, including studying variants of this baseline algorithm that optimize the subset chosen using different techniques and criteria. We find that the baseline algorithm remains competitive with its variants in most settings, and build a better understanding of the practical importance of a bias-variance tradeoff inherent in solutions to the contribution bounding problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef51b72d-6195-4b8b-bd23-0ecaffc21231Builds on5
- Practical and Private (Deep) Learning Without Sampling or ShufflingPeter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar et al.ICML 2021 · 239 citations
- Improved Differential Privacy for SGD via Optimal Private Linear Operators on Adaptive StreamsSergey Denisov, H. Brendan McMahan, John Rush, Adam D. Smith et al.NeurIPS 2022 · 96 citations
- Multi-Epoch Matrix Factorization Mechanisms for Private Machine LearningChristopher A. Choquette-Choo, Hugh Brendan McMahan, J. Keith Rush, Abhradeep Guha ThakurtaICML 2023 · 62 citations
- Shifted Inverse: A General Mechanism for Monotonic Functions under User Differential PrivacyJuanru Fang, Wei Dong, Ke YiCCS 2022 · 13 citations
- Time-Aware Projections: Truly Node-Private Graph Statistics under Continual ObservationPalak Jain, Adam Smith, Connor WagamanS&P 2024 · 11 citations
Related papers
- Smoothly Bounding User Contributions in Differential PrivacyAlessandro Epasto, Mohammad Mahdian, Jieming Mao, Vahab S. Mirrokni et al.NeurIPS 2020 · 17 citations
- Algorithms for bounding contribution for histogram estimation under user-level privacyYuhan Liu, Ananda Theertha Suresh, Wennan Zhu, Peter Kairouz et al.ICML 2023 · 14 citations
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale et al.NeurIPS 2021 · 113 citations
- User-Level Differential Privacy With Few Examples Per UserBadih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi et al.NeurIPS 2023 · 19 citations
- Task-aware Privacy Preservation for Multi-dimensional DataJiangnan Cheng, Ao Tang, Sandeep ChinchaliICML 2022 · 8 citations
