Rethinking the Security of DP-SGD: A Corrected Analysis of Differentially Private Machine Learning
Wenhao Wang, Shujie Cui, Hui Cui, Xingliang Yuan
摘要
Differentially Private Stochastic Gradient Descent (DP-SGD) has been widely adopted to protect training data in machine learning. The privacy guarantee of DP-SGD and the DP-mechanisms built upon it is usually analyzed through a formal security game, in which an adversary infers whether a particular individual data record is included in the training dataset based on the mechanism's output. Privacy leakage is characterized by the adversary's privacy curve, which reports the false negative rate (FNR) as a function of the false positive rate (FPR). The privacy guarantee is defined as the lower bound of this curve.
We observe that the privacy guarantees claimed for these mechanisms in many existing papers are derived from a mismatched game setting.
Specifically, they formalize the mechanisms as the Subsampled Gaussian Mechanism (SGM), where Gaussian noise is added to the sum of gradients computed from a Poisson-sampled batch of data. Indeed, the training procedure in these mechanisms introduces an additional normalization step: the noisy sum is further normalized either by the expected batch size or by the sampled batch size. Thus, these mechanisms should be formalized as either the Expected-Averaged SGM (EASGM) or the Batch-Averaged SGM (ASGM). The privacy auditing of DP-SGD suffers from the same issue, as it assesses the privacy guarantee of DP-SGD by treating it as an SGM.
We therefore re-analyze the privacy guarantee of these mechanisms under the corresponding EASGM and ASGM formalizations. Our analysis shows that, in theory, these DP-SGD mechanisms can yield weaker privacy guarantees than the SGM-based guarantee, suggesting that, in some settings, the true privacy leakage can exceed the reported SGM-based guarantee.
We also empirically audit the leakage of implementations of four state-of-the-art DP-SGD algorithms, including the implementation used in Meta's Opacus library, and show that empirical leakage exceeds the SGM-based guarantees. Finally, we conduct a thorough code audit of Opacus versions v0.9.0-v1.5.4 and derive a privacy guarantee for the latest Opacus implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 被引用 325 次
- Adversary Instantiation: Lower Bounds for Differentially Private Machine LearningMilad Nasr, Shuang Song, Abhradeep Thakurta, Nicolas Papernot 等S&P 2021 · 被引用 288 次
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 被引用 259 次
- Privacy Auditing with One (1) Training RunThomas Steinke, Milad Nasr, Matthew JagielskiNeurIPS 2023 · 被引用 178 次
相关 Paper
- To Shuffle or not to Shuffle: Auditing DP-SGD with ShufflingMeenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De CristofaroNDSS 2026 · 被引用 11 次
- How Private are DP-SGD Implementations?Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar 等ICML 2024 · 被引用 25 次
- Auditing Differentially Private Machine Learning: How Private is Private SGD?Matthew Jagielski, Jonathan R. Ullman, Alina OpreaNeurIPS 2020 · 被引用 354 次
- Tighter Privacy Auditing of DP-SGD in the Hidden State Threat ModelTudor Ioan Cebere, Aurélien Bellet, Nicolas PapernotICLR 2025 · 被引用 1 次
- Scalable DP-SGD: Shuffling vs. Poisson SubsamplingLynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar 等NeurIPS 2024 · 被引用 29 次
