Public-data Assisted Private Stochastic Optimization: Power and Limitations
Enayat Ullah, Michael Menart, Raef Bassily, Cristóbal Guzmán, Raman Arora
Abstract
We study the limits and capability of public-data assisted differentially private (PA-DP) algorithms. Specifically, we focus on the problem of stochastic convex optimization (SCO) with either labeled or unlabeled public data. For complete/labeled public data, we show that any -PA-DP has excess risk , where is the dimension, is the number of public samples, is the number of private samples, and . These lower bounds are established via our new lower bounds for PA-DP mean estimation, which are of a similar form. Up to constant factors, these lower bounds show that the simple strategy of either treating all data as private or discarding the private data, is optimal. We also study PA-DP supervised learning with unlabeled public samples. In contrast to our previous result, we here show novel methods for leveraging public data in private supervised learning. For generalized linear models (GLM) with unlabeled public data, we show an efficient algorithm which, given unlabeled public samples, achieves the dimension independent rate . We develop new lower bounds for this setting which shows that this rate cannot be improved with more public samples, and any fewer public samples leads to a worse rate. Finally, we provide extensions of this result to general hypothesis classes with finite fat-shattering dimension with applications to neural networks and non-Euclidean geometries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b74bef5-87a4-4d8e-82ee-2bc7d1fc4db9Cited by top-tier papers2
- Optimal Differentially Private Model Training with Public DataAndrew Lowy, Zeman Li, Tianjian Huang, Meisam RazaviyaynICML 2024 · 9 citations
- On the Benefits of Public Representations for Private Transfer Learning under Distribution ShiftPratiksha Thaker, Amrith Setlur, Steven Z. Wu, Virginia SmithNeurIPS 2024 · 7 citations
Builds on12
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- Do not Let Privacy Overbill Utility: Gradient Embedding Perturbation for Private LearningDa Yu, Huishuai Zhang, Wei Chen, Tie-Yan LiuICLR 2021 · 133 citations
- Bypassing the Ambient Dimension: Private SGD with Gradient Subspace IdentificationYingxue Zhou, Steven Wu, Arindam BanerjeeICLR 2021 · 118 citations
- Differentially Private Stochastic Optimization: New Results in Convex and Non-Convex SettingsRaef Bassily, Cristóbal Guzmán, Michael MenartNeurIPS 2021 · 68 citations
Related papers
- On Traceability in ℓp Stochastic Convex OptimizationSasha Voitovych, Mahdi Haghifam, Idan Attias, Gintare Karolina Dziugaite et al.NeurIPS 2025
- Faster Algorithms for User-Level Private Stochastic Convex OptimizationAndrew Lowy, Daogao Liu, Hilal AsiNeurIPS 2024 · 4 citations
- Oracle-Efficient Differentially Private Learning with Public DataAdam Block, Mark Bun, Rathin Desai, Abhishek Shetty et al.NeurIPS 2024 · 6 citations
- Differentially Private Worst-group Risk MinimizationXinyu Zhou, Raef BassilyICML 2024 · 7 citations
- User-level Private Stochastic Convex Optimization with Optimal RatesRaef Bassily, Ziteng SunICML 2023 · 17 citations
