Scalable DP-SGD: Shuffling vs. Poisson Subsampling
Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi, Amer Sinha, Chiyuan Zhang
摘要
We provide new lower bounds on the privacy guarantee of the multi-epoch Adaptive Batch Linear Queries (ABLQ) mechanism with shuffled batch sampling, demonstrating substantial gaps when compared to Poisson subsampling; prior analysis was limited to a single epoch. Since the privacy analysis of Differentially Private Stochastic Gradient Descent (DP-SGD) is obtained by analyzing the ABLQ mechanism, this brings into serious question the common practice of implementing shuffling-based DP-SGD, but reporting privacy parameters as if Poisson subsampling was used. To understand the impact of this gap on the utility of trained machine learning models, we introduce a practical approach to implement Poisson subsampling at scale using massively parallel computation, and efficiently train models with the same. We compare the utility of models trained with Poisson-subsampling-based DP-SGD, and the optimistic estimates of utility when using shuffling, via our new lower bounds on the privacy guarantee of ABLQ with shuffling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Privacy amplification by random allocationMoshe Shenfeld, Vitaly FeldmanNeurIPS 2025 · 被引用 18 次
- To Shuffle or not to Shuffle: Auditing DP-SGD with ShufflingMeenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De CristofaroNDSS 2026 · 被引用 11 次
- Efficient privacy loss accounting for subsampling and random allocationVitaly Feldman, Moshe ShenfeldICML 2026 · 被引用 5 次
- Fundamental Limitations of Favorable Privacy–Utility Guarantees for DP-SGDMurat Bilgehan Ertan), Marten van Dijk)CCS 2026 · 被引用 3 次
- Convex Approximation of Two-Layer ReLU Networks for Hidden State Differential PrivacyRob Romijnders, Antti KoskelaNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper11
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 被引用 325 次
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 被引用 259 次
- Practical and Private (Deep) Learning Without Sampling or ShufflingPeter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar 等ICML 2021 · 被引用 239 次
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private GeneratorsDingfan Chen, Tribhuvanesh Orekondy, Mario FritzNeurIPS 2020 · 被引用 228 次
相关 Paper
- How Private are DP-SGD Implementations?Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar 等ICML 2024 · 被引用 25 次
- Rethinking the Security of DP-SGD: A Corrected Analysis of Differentially Private Machine LearningWenhao Wang, Shujie Cui, Hui Cui, Xingliang YuanCCS 2026
- Near-Exact Privacy Amplification for Matrix MechanismsChristopher A. Choquette-Choo, Arun Ganesh, Saminul Haque, Thomas Steinke 等ICLR 2025
- Differentially Private Stochastic Gradient Descent with Fixed-Size Minibatches: Tighter RDP Guarantees with or without ReplacementJeremiah Birrell, Reza Ebrahimi, Rouzbeh Behnia, Jason PachecoNeurIPS 2024 · 被引用 10 次
- Auditing Differentially Private Machine Learning: How Private is Private SGD?Matthew Jagielski, Jonathan R. Ullman, Alina OpreaNeurIPS 2020 · 被引用 354 次
