Effectively Using Public Data in Privacy Preserving Machine Learning
Milad Nasr, Saeed Mahloujifar, Xinyu Tang, Prateek Mittal, Amir Houmansadr
摘要
Differentially private (DP) machine learning techniques are notorious for their degradation of model utility (e.g., they degrade classification accuracy). A recent line of work has demonstrated that leveraging public data can improve the trade-off between privacy and utility when training models with DP guarantee. In this work, we further explore the potential of using public data in DP models, showing that utility gains can in fact be significantly higher than what shown in prior works. Specifically, we introduce DOPE-SGD, a modified DP-SGD algorithm that leverages public data during its training. DOPE-SGD uses public data in two complementary ways: (1) it uses advance augmentation techniques that leverages public data to generate synthetic data that is effectively embedded in multiple steps of the training pipeline; (2) it uses a modified gradient clipping mechanism in DP-SGD to change the origin of gradient vectors using the information inferred from available public data, therefore boosting utility. We also introduce a technique to ensemble intermediate DP models by leveraging the post processing property of differential privacy to further improve the accuracy of the predictions. Our experimental results demonstrate the effectiveness of our approach in improving the state-of-the-art in DP machine learning across multiple datasets, network architectures, and application domains. For instance, assuming access to 2, 000 public images, and for a privacy budget of ε = 2, δ = 10 -5 , our technique achieves an accuracy of 75.1% on CIFAR10, significantly higher than 68.1% achieved by the state of the art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Why Is Public Pretraining Necessary for Private Model Training?Arun Ganesh, Mahdi Haghifam, Milad Nasr, Sewoong Oh 等ICML 2023 · 被引用 47 次
- Differentially Private Image Classification by Learning Priors from Random ProcessesXinyu Tang, Ashwinee Panda, Vikash Sehwag, Prateek MittalNeurIPS 2023 · 被引用 34 次
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway 等ICML 2024 · 被引用 30 次
- A New Linear Scaling Rule for Private Adaptive Hyperparameter OptimizationAshwinee Panda, Xinyu Tang, Saeed Mahloujifar, Vikash Sehwag 等ICML 2024 · 被引用 15 次
- Public-data Assisted Private Stochastic Optimization: Power and LimitationsEnayat Ullah, Michael Menart, Raef Bassily, Cristóbal Guzmán 等NeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper11
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 被引用 543 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
相关 Paper
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao 等USENIX Security 2024 · 被引用 23 次
- In-Distribution Public Data Synthesis With Diffusion Models for Differentially Private Image ClassificationJinseong Park, Yujin Choi, Jaewook LeeCVPR 2024
- Public Data-Assisted Mirror Descent for Private Model TrainingEhsan Amid, Arun Ganesh, Rajiv Mathews, Swaroop Ramaswamy 等ICML 2022 · 被引用 61 次
- Machine Learning with Privacy for Protected AttributesSaeed Mahloujifar, Chuan Guo, G. Edward Suh, Kamalika ChaudhuriS&P 2025
- Differentially Private Prototypes for Imbalanced Transfer LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischAAAI 2025 · 被引用 4 次
