On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
Aki Rehn, Linzh Zhao, Mikko A. Heikkilä, Antti Honkela
摘要
Differentially private (DP) transfer learning, i.e., fine-tuning a pretrained model on private data, is the current state-of-the-art approach for training large models under privacy constraints. We focus on two key hyperparameters in this setting: the clipping bound and batch size . We show a clear mismatch between the current theoretical understanding of how to choose an optimal (stronger privacy requires smaller ) and empirical outcomes (larger performs better under strong privacy), caused by changes in the gradient distributions. Assuming a limited compute budget (fixed epochs), we demonstrate that the existing heuristics for tuning do not work, while cumulative DP noise better explains whether smaller or larger batches perform better. We also highlight how the common practice of using a single setting across tasks can lead to suboptimal performance. We find that performance drops especially when moving between loose and tight privacy and between plentiful and limited compute, which we explain by analyzing clipping as a form of gradient re-weighting and examining cumulative DP noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 被引用 425 次
相关 Paper
- Adaptive Methods Are Preferable in High Privacy Settings: An SDE PerspectiveEnea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov, Aurélien Lucchi 等ICLR 2026 · 被引用 2 次
- TAN Without a Burn: Scaling Laws of DP-SGDTom Sander, Pierre Stock, Alexandre SablayrollesICML 2023 · 被引用 58 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
- Online Sensitivity Optimization in Differentially Private LearningFilippo Galli, Catuscia Palamidessi, Tommaso CucinottaAAAI 2024 · 被引用 3 次
- The Role of Adaptive Optimizers for Honest Private Hyperparameter SelectionShubhankar Mohapatra, Sajin Sasy, Xi He, Gautam Kamath 等AAAI 2022 · 被引用 35 次
