Differentially Private Optimization on Large Model at Small Cost
Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, George Karypis
摘要
Differentially private (DP) optimization is the standard paradigm to learn large neural networks that are accurate and privacy-preserving. The computational cost for DP deep learning, however, is notoriously heavy due to the per-sample gradient clipping. Existing DP implementations are 2-1000X more costly in time and space complexity than the standard (non-private) training. In this work, we develop a novel Book-Keeping (BK) technique that implements existing DP optimizers (thus achieving the same accuracy), with a substantial improvement on the computational cost. Specifically, BK enables DP training on large models and high dimensional data to be roughly as fast and memory-saving as the standard training, whereas previous DP algorithms can be inefficient or incapable of training due to memory error. The computational advantage of BK is supported by the complexity analysis as well as extensive experiments on vision and language tasks. Our implementation achieves state-of-the-art (SOTA) accuracy with very small extra cost: on GPT2 and at almost the same memory cost (<1% overhead), BK has 1.03X the time complexity of the standard training (0.83X training speed in practice), and 0.61X the time complexity of the most efficient DP implementation (1.36X training speed in practice). We open-source the codebase for the BK algorithm at the FastDP library (https://github.com/awslabs/fast-differential-privacy).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal 等NeurIPS 2024 · 被引用 91 次
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan 等VLDB 2024 · 被引用 66 次
- Differentially Private Bias-Term Fine-tuning of Foundation ModelsZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisICML 2024 · 被引用 59 次
- Privacy-Preserving In-Context Learning for Large Language ModelsTong Wu, Ashwinee Panda, Jiachen T. Wang, Prateek MittalICLR 2024 · 被引用 58 次
- Why Is Public Pretraining Necessary for Private Model Training?Arun Ganesh, Mahdi Haghifam, Milad Nasr, Sewoong Oh 等ICML 2023 · 被引用 47 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Exploring the Limits of Differentially Private Deep Learning with Group-wise ClippingJiyan He, Xuechen Li, Da Yu, Huishuai Zhang 等ICLR 2023 · 被引用 4 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
- Scalable and Efficient Training of Large Convolutional Neural Networks with Differential PrivacyZhiqi Bu, Jialin Mao, Shiyun XuNeurIPS 2022 · 被引用 70 次
- Towards hyperparameter-free optimization with differential privacyRuixuan Liu, Zhiqi BuICLR 2025
- Fast and Memory Efficient Differentially Private-SGD via JL ProjectionsZhiqi Bu, Sivakanth Gopi, Janardhan Kulkarni, Yin Tat Lee 等NeurIPS 2021 · 被引用 49 次
