Rethinking the Security of DP-SGD: A Corrected Analysis of Differentially Private Machine Learning
Wenhao Wang, Shujie Cui, Hui Cui, Xingliang Yuan
Abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) has been widely adopted to protect training data in machine learning. The privacy guarantee of DP-SGD and the DP-mechanisms built upon it is usually analyzed through a formal security game, in which an adversary infers whether a particular individual data record is included in the training dataset based on the mechanism's output. Privacy leakage is characterized by the adversary's privacy curve, which reports the false negative rate (FNR) as a function of the false positive rate (FPR). The privacy guarantee is defined as the lower bound of this curve.
We observe that the privacy guarantees claimed for these mechanisms in many existing papers are derived from a mismatched game setting.
Specifically, they formalize the mechanisms as the Subsampled Gaussian Mechanism (SGM), where Gaussian noise is added to the sum of gradients computed from a Poisson-sampled batch of data. Indeed, the training procedure in these mechanisms introduces an additional normalization step: the noisy sum is further normalized either by the expected batch size or by the sampled batch size. Thus, these mechanisms should be formalized as either the Expected-Averaged SGM (EASGM) or the Batch-Averaged SGM (ASGM). The privacy auditing of DP-SGD suffers from the same issue, as it assesses the privacy guarantee of DP-SGD by treating it as an SGM.
We therefore re-analyze the privacy guarantee of these mechanisms under the corresponding EASGM and ASGM formalizations. Our analysis shows that, in theory, these DP-SGD mechanisms can yield weaker privacy guarantees than the SGM-based guarantee, suggesting that, in some settings, the true privacy leakage can exceed the reported SGM-based guarantee.
We also empirically audit the leakage of implementations of four state-of-the-art DP-SGD algorithms, including the implementation used in Meta's Opacus library, and show that empirical leakage exceeds the SGM-based guarantees. Finally, we conduct a thorough code audit of Opacus versions v0.9.0-v1.5.4 and derive a privacy guarantee for the latest Opacus implementation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f075d4f9-43a5-48f4-b6f9-e545b1913a27Builds on22
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 325 citations
- Adversary Instantiation: Lower Bounds for Differentially Private Machine LearningMilad Nasr, Shuang Song, Abhradeep Thakurta, Nicolas Papernot et al.S&P 2021 · 288 citations
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 259 citations
- Privacy Auditing with One (1) Training RunThomas Steinke, Milad Nasr, Matthew JagielskiNeurIPS 2023 · 178 citations
Related papers
- To Shuffle or not to Shuffle: Auditing DP-SGD with ShufflingMeenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De CristofaroNDSS 2026 · 11 citations
- How Private are DP-SGD Implementations?Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar et al.ICML 2024 · 25 citations
- Auditing Differentially Private Machine Learning: How Private is Private SGD?Matthew Jagielski, Jonathan R. Ullman, Alina OpreaNeurIPS 2020 · 354 citations
- Tighter Privacy Auditing of DP-SGD in the Hidden State Threat ModelTudor Ioan Cebere, Aurélien Bellet, Nicolas PapernotICLR 2025 · 1 citation
- Scalable DP-SGD: Shuffling vs. Poisson SubsamplingLynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar et al.NeurIPS 2024 · 29 citations
