Public Data-Assisted Mirror Descent for Private Model Training
Ehsan Amid, Arun Ganesh, Rajiv Mathews, Swaroop Ramaswamy, Shuang Song, Thomas Steinke, Vinith M. Suriyakumar, Om Thakkar, Abhradeep Thakurta
Abstract
In this paper, we revisit the problem of using in-distribution public data to improve the privacy/utility trade-offs for differentially private (DP) model training. (Here, public data refers to auxiliary data sets that have no privacy concerns.) We design a natural variant of DP mirror descent, where the DP gradients of the private/sensitive data act as the linear term, and the loss generated by the public data as the mirror map. We show that, for linear regression with feature vectors drawn from a non-isotropic sub-Gaussian distribution, our algorithm, PDA-DPMD (a variant of mirror descent), provides population risk guarantees that are asymptotically better than the best known guarantees under DP (without having access to public data), when the number of public data samples () is sufficiently large. We further show that our algorithm has natural"noise stability"properties that control the variance due to noise added to ensure DP. We demonstrate the efficacy of our algorithm by showing privacy/utility trade-offs on four benchmark datasets (StackOverflow, WikiText-2, CIFAR-10, and EMNIST). We show that our algorithm not only significantly improves over traditional DP-SGD, which does not have access to public data, but to our knowledge is the first to improve over DP-SGD on models that have been pre-trained with public data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6cd0b87-00cc-44c0-8d5c-df0bc86276faCited by top-tier papers25
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- When Does Differentially Private Learning Not Suffer in High Dimensions?Xuechen Li, Daogao Liu, Tatsunori B. Hashimoto, Huseyin A. Inan et al.NeurIPS 2022 · 64 citations
- Why Is Public Pretraining Necessary for Private Model Training?Arun Ganesh, Mahdi Haghifam, Milad Nasr, Sewoong Oh et al.ICML 2023 · 47 citations
- Private Adaptive Optimization with Side informationTian Li, Manzil Zaheer, Sashank J. Reddi, Virginia SmithICML 2022 · 46 citations
- Private Estimation with Public DataAlex Bie, Gautam Kamath, Vikrant SinghalNeurIPS 2022 · 40 citations
Builds on16
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 325 citations
- Stability of Stochastic Gradient Descent on Nonsmooth Convex LossesRaef Bassily, Vitaly Feldman, Cristóbal Guzmán, Kunal TalwarNeurIPS 2020 · 240 citations
- Practical and Private (Deep) Learning Without Sampling or ShufflingPeter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar et al.ICML 2021 · 239 citations
Related papers
- Optimal Differentially Private Model Training with Public DataAndrew Lowy, Zeman Li, Tianjian Huang, Meisam RazaviyaynICML 2024 · 9 citations
- Effectively Using Public Data in Privacy Preserving Machine LearningMilad Nasr, Saeed Mahloujifar, Xinyu Tang, Prateek Mittal et al.ICML 2023 · 22 citations
- Differentially Private Image Classification by Learning Priors from Random ProcessesXinyu Tang, Ashwinee Panda, Vikash Sehwag, Prateek MittalNeurIPS 2023 · 34 citations
- Differentially Private Prototypes for Imbalanced Transfer LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischAAAI 2025 · 4 citations
- Have it your way: Individualized Privacy Assignment for DP-SGDFranziska Boenisch, Christopher Mühl, Adam Dziedzic, Roy Rinberg et al.NeurIPS 2023 · 39 citations
