Optimal Differentially Private Model Training with Public Data
Andrew Lowy, Zeman Li, Tianjian Huang, Meisam Razaviyayn
Abstract
Differential privacy (DP) ensures that training a machine learning model does not leak private data. In practice, we may have access to auxiliary public data that is free of privacy concerns. In this work, we assume access to a given amount of public data and settle the following fundamental open questions: 1. What is the optimal (worst-case) error of a DP model trained over a private data set while having access to side public data? 2. How can we harness public data to improve DP model training in practice? We consider these questions in both the local and central models of pure and approximate DP. To answer the first question, we prove tight (up to log factors) lower and upper bounds that characterize the optimal error rates of three fundamental problems: mean estimation, empirical risk minimization, and stochastic convex optimization. We show that the optimal error rates can be attained (up to log factors) by either discarding private data and training a public model, or treating public data like it is private and using an optimal DP algorithm. To address the second question, we develop novel algorithms that are "even more optimal" (i.e. better constants) than the asymptotically optimal approaches described above. For local DP mean estimation, our algorithm is optimal including constants. Empirically, our algorithms show benefits over the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- How to Make the Gradients Small Privately: Improved Rates for Differentially Private Non-Convex OptimizationAndrew Lowy, Jonathan R. Ullman, Stephen J. WrightICML 2024 · 11 citations
- On the Benefits of Public Representations for Private Transfer Learning under Distribution ShiftPratiksha Thaker, Amrith Setlur, Steven Z. Wu, Virginia SmithNeurIPS 2024 · 7 citations
- Public-data Assisted Private Stochastic Optimization: Power and LimitationsEnayat Ullah, Michael Menart, Raef Bassily, Cristóbal Guzmán et al.NeurIPS 2024 · 6 citations
- Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public DataAhmed Mehdi Inane, Vincent Quirion, Gintare Karolina Dziugaite, Ioannis MitliagkasICML 2026
Builds on22
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
Related papers
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale et al.NeurIPS 2021 · 113 citations
- Public Data-Assisted Mirror Descent for Private Model TrainingEhsan Amid, Arun Ganesh, Rajiv Mathews, Swaroop Ramaswamy et al.ICML 2022 · 61 citations
- Differentially Private Learning with Small Public DataJun Wang, Zhi-Hua ZhouAAAI 2020 · 25 citations
- User-level Private Stochastic Convex Optimization with Optimal RatesRaef Bassily, Ziteng SunICML 2023 · 17 citations
- Effectively Using Public Data in Privacy Preserving Machine LearningMilad Nasr, Saeed Mahloujifar, Xinyu Tang, Prateek Mittal et al.ICML 2023 · 22 citations
