Subsampling is not Magic: Why Large Batch Sizes Work for Differentially Private Stochastic Optimisation
Ossi Räisä, Joonas Jälkö, Antti Honkela
Abstract
We study how the batch size affects the total gradient variance in differentially private stochastic gradient descent (DP-SGD), seeking a theoretical explanation for the usefulness of large batch sizes. As DP-SGD is the basis of modern DP deep learning, its properties have been widely studied, and recent works have empirically found large batch sizes to be beneficial. However, theoretical explanations of this benefit are currently heuristic at best. We first observe that the total gradient variance in DP-SGD can be decomposed into subsampling-induced and noise-induced variances. We then prove that in the limit of an infinite number of iterations, the effective noise-induced variance is invariant to the batch size. The remaining subsampling-induced variance decreases with larger batch sizes, so large batches reduce the effective total gradient variance. We confirm numerically that the asymptotic regime is relevant in practical settings when the batch size is not small, and find that outside the asymptotic regime, the total gradient variance decreases even more with large batch sizes. We also find a sufficient condition that implies that large batch sizes similarly reduce effective DP noise variance for one iteration of DP-SGD 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5121f5ce-2669-4928-9953-8aba22f23e77Cited by top-tier papers2
- On Optimal Hyperparameters for Differentially Private Deep Transfer LearningAki Rehn, Linzh Zhao, Mikko A. Heikkilä, Antti HonkelaICLR 2026 · 2 citations
- The Adverse Effects of Omitting Records in Differential Privacy: How Sampling and Suppression Degrade the Privacy–Utility TradeoffÀlex Miranda-Pascual, Javier Parra-Arnau, Thorsten StrufeUSENIX Security 2026
Builds on5
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 259 citations
- Automatic Clipping: Differentially Private Deep Learning Made Easier and StrongerZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisNeurIPS 2023 · 140 citations
- The Saddle-Point Method in Differential PrivacyWael Alghamdi, Juan Felipe Gómez, Shahab Asoodeh, Flávio P. Calmon et al.ICML 2023 · 16 citations
Related papers
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 122 citations
- Bypassing the Ambient Dimension: Private SGD with Gradient Subspace IdentificationYingxue Zhou, Steven Wu, Arindam BanerjeeICLR 2021 · 118 citations
- A Theory to Instruct Differentially-Private Learning via Clipping Bias ReductionHanshen Xiao, Zihang Xiang, Di Wang, Srinivas DevadasS&P 2023
- DP-PCA: Statistically Optimal and Differentially Private PCAXiyang Liu, Weihao Kong, Prateek Jain, Sewoong OhNeurIPS 2022 · 38 citations
- Enhancing DPSGD via Per-Sample Momentum and Low-Pass FilteringXincheng Xu, Thilina Ranbaduge, Qing Wang, Thierry Rakotoarivelo et al.AAAI 2026
