Pre-training Differentially Private Models with Limited Public Data
Zhiqi Bu, Xinwei Zhang, Sheng Zha, Mingyi Hong, George Karypis
Abstract
The superior performance of large foundation models relies on the use of massive amounts of high-quality data, which often contain sensitive, private and copyrighted material that requires formal protection. While differential privacy (DP) is a prominent method to gauge the degree of security provided to the models, its application is commonly limited to the model fine-tuning stage, due to the performance degradation when applying DP during the pre-training stage. Consequently, DP is yet not capable of protecting a substantial portion of the data used during the initial pre-training process. In this work, we first provide a theoretical understanding of the efficacy of DP training by analyzing the per-iteration loss improvement. We make a key observation that DP optimizers' performance degradation can be significantly mitigated by the use of limited public data, which leads to a novel DP continual pre-training strategy. Empirically, using only 10% of public data, our strategy can achieve DP accuracy of 41.5% on ImageNet-21k (with ), as well as non-DP accuracy of 55.7% and and 60.0% on downstream tasks Places365 and iNaturalist-2021, respectively, on par with state-of-the-art standard pre-training and substantially outperforming existing DP pre-trained models. Our DP pre-trained models are released in fastDP library (https://github.com/awslabs/fast-differential-privacy/releases/tag/v2.1)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Understanding Private Learning From Feature PerspectiveMeng Ding, Mingxi Lei, Shaopeng Fu, Shaowei Wang et al.ICML 2026 · 2 citations
- FIBER: A Differentially Private Optimizer with Filter-Aware Innovation Bias CorrectionMINH DUC DO, Thao Do, Minh Hoang, Anh Le Duc Tran et al.ICML 2026 · 1 citation
- MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic ApproximationLu Li, Tianyu Zhang, Zhiqi Bu, Suyuchen Wang et al.ICLR 2025
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Differentially Private Bias-Term Fine-tuning of Foundation ModelsZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisICML 2024 · 59 citations
- DOPPLER: Differentially Private Optimizers with Low-pass Filter for Privacy Noise ReductionXinwei Zhang, Zhiqi Bu, Mingyi Hong, Meisam RazaviyaynNeurIPS 2024 · 10 citations
- Differentially Private Optimization on Large Model at Small CostZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisICML 2023 · 85 citations
- In-Distribution Public Data Synthesis With Diffusion Models for Differentially Private Image ClassificationJinseong Park, Yujin Choi, Jaewook LeeCVPR 2024
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao et al.USENIX Security 2024 · 23 citations
