To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling
Meenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De Cristofaro
摘要
The Differentially Private Stochastic Gradient Descent (DP-SGD) algorithm supports the training of machine learning (ML) models with formal Differential Privacy (DP) guarantees. Traditionally, DP-SGD processes training data in batches using Poisson subsampling to select each batch at every iteration. More recently, shuffling has become a common alternative due to its better compatibility and lower computational overhead. However, computing tight theoretical DP guarantees under shuffling remains an open problem. As a result, models trained with shuffling are often evaluated as if Poisson subsampling were used, which might result in incorrect privacy guarantees. This raises a compelling research question: can we verify whether there are gaps between the theoretical DP guarantees reported by state-of-the-art models using shuffling and their actual leakage? To do so, we define novel DP-auditing procedures to analyze DP-SGD with shuffling and measure their ability to tightly estimate privacy leakage vis-`a-vis batch sizes, privacy budgets, and threat models. Overall, we demonstrate that DP models trained using this approach have considerably overestimated their privacy guarantees (by up to 4 times). However, we also find that the gap between the theoretical Poisson DP guarantees and the actual privacy leakage from shuffling is not uniform across all parameter settings and threat models. Finally, we study two common variations of the shuffling procedure that result in even further privacy leakage (up to 10 times). Overall, our work highlights the risk of using shuffling instead of Poisson subsampling in the absence of rigorous analysis methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Privacy amplification by random allocationMoshe Shenfeld, Vitaly FeldmanNeurIPS 2025 · 被引用 18 次
- Efficient privacy loss accounting for subsampling and random allocationVitaly Feldman, Moshe ShenfeldICML 2026 · 被引用 5 次
- General-Purpose f-DP Estimation and Auditing in a Black-Box SettingÖnder Askin, Holger Dette, Martin Dunsche, Tim Kutta 等USENIX Security 2025
- Curation Leaks: Membership Inference Attacks against Data Curation for Machine LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischICLR 2026
- Rethinking the Security of DP-SGD: A Corrected Analysis of Differentially Private Machine LearningWenhao Wang, Shujie Cui, Hui Cui, Xingliang YuanCCS 2026
它引用的顶会 Paper30
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 被引用 586 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
相关 Paper
- How Private are DP-SGD Implementations?Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar 等ICML 2024 · 被引用 25 次
- Scalable DP-SGD: Shuffling vs. Poisson SubsamplingLynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar 等NeurIPS 2024 · 被引用 29 次
- Tighter Privacy Auditing of DP-SGD in the Hidden State Threat ModelTudor Ioan Cebere, Aurélien Bellet, Nicolas PapernotICLR 2025 · 被引用 1 次
- Tight Auditing of Differentially Private Machine LearningMilad Nasr, Jamie Hayes, Thomas Steinke, Borja Balle 等USENIX Security 2023
- Nearly Tight Black-Box Auditing of Differentially Private Machine LearningMeenatchi Sundaram Muthu Selva Annamalai, Emiliano De CristofaroNeurIPS 2024 · 被引用 32 次
