On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
Aki Rehn, Linzh Zhao, Mikko A. Heikkilä, Antti Honkela
Abstract
Differentially private (DP) transfer learning, i.e., fine-tuning a pretrained model on private data, is the current state-of-the-art approach for training large models under privacy constraints. We focus on two key hyperparameters in this setting: the clipping bound and batch size . We show a clear mismatch between the current theoretical understanding of how to choose an optimal (stronger privacy requires smaller ) and empirical outcomes (larger performs better under strong privacy), caused by changes in the gradient distributions. Assuming a limited compute budget (fixed epochs), we demonstrate that the existing heuristics for tuning do not work, while cumulative DP noise better explains whether smaller or larger batches perform better. We also highlight how the common practice of using a single setting across tasks can lead to suboptimal performance. We find that performance drops especially when moving between loose and tight privacy and between plentiful and limited compute, which we explain by analyzing clipping as a form of gradient re-weighting and examining cumulative DP noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e484ce8-6045-435d-92bc-332b4e65e9e6Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 425 citations
Related papers
- Adaptive Methods Are Preferable in High Privacy Settings: An SDE PerspectiveEnea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov, Aurélien Lucchi et al.ICLR 2026 · 2 citations
- TAN Without a Burn: Scaling Laws of DP-SGDTom Sander, Pierre Stock, Alexandre SablayrollesICML 2023 · 58 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Online Sensitivity Optimization in Differentially Private LearningFilippo Galli, Catuscia Palamidessi, Tommaso CucinottaAAAI 2024 · 3 citations
- The Role of Adaptive Optimizers for Honest Private Hyperparameter SelectionShubhankar Mohapatra, Sajin Sasy, Xi He, Gautam Kamath et al.AAAI 2022 · 35 citations
