Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
Anvith Thudi, Hengrui Jia, Casey Meehan, Ilia Shumailov, Nicolas Papernot
摘要
Differentially private stochastic gradient descent (DP-SGD) is the canonical approach to private deep learning. While the current privacy analysis of DP-SGD is known to be tight in some settings, several empirical results suggest that models trained on common benchmark datasets leak significantly less privacy for many datapoints. Yet, despite past attempts, a rigorous explanation for why this is the case has not been reached. Is it because there exist tighter privacy upper bounds when restricted to these dataset settings, or are our attacks not strong enough for certain datapoints? In this paper, we provide the first per-instance (i.e., ``data-dependent") DP analysis of DP-SGD. Our analysis captures the intuition that points with similar neighbors in the dataset enjoy better data-dependent privacy than outliers. Formally, this is done by modifying the per-step privacy analysis of DP-SGD to introduce a dependence on the distribution of model updates computed from a training dataset. We further develop a new composition theorem to effectively use this new per-step analysis to reason about an entire training run. Put all together, our evaluation shows that this novel DP-SGD analysis allows us to now formally show that DP-SGD leaks significantly less privacy for many datapoints (when trained on common benchmarks) than the current data-independent guarantee. This implies privacy attacks will necessarily fail against many datapoints if the adversary does not have sufficient control over the possible training datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Graphical vs. Deep Generative Models: Measuring the Impact of Differentially Private Mechanisms and Budgets on UtilityGeorgi Ganev, Kai Xu, Emiliano De CristofaroCCS 2024 · 被引用 5 次
- Efficient Public Verification of Private ML via RegularizationZoë R Bell, Anvith Thudi, Olive Franzese-McLaughlin, Nicolas Papernot 等ICML 2026 · 被引用 1 次
- Trustworthy Machine Learning through Data-Specific IndistinguishabilityHanshen Xiao, Zhen Yang, G. Edward SuhICML 2025
- Free Record-Level Privacy Risk Evaluation Through Artifact-Based MethodsJoseph Pollock, Igor Shilov, Euodia Dodd, Yves-Alexandre de MontjoyeUSENIX Security 2025
- GPM: The Gaussian Pancake Mechanism for Planting Undetectable Backdoors in Differential PrivacyHaochen Sun, Xi HeSIGMOD 2026
它引用的顶会 Paper21
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Auditing Differentially Private Machine Learning: How Private is Private SGD?Matthew Jagielski, Jonathan R. Ullman, Alina OpreaNeurIPS 2020 · 被引用 354 次
相关 Paper
- Adversary Instantiation: Lower Bounds for Differentially Private Machine LearningMilad Nasr, Shuang Song, Abhradeep Thakurta, Nicolas Papernot 等S&P 2021 · 被引用 288 次
- Bounding training data reconstruction in DP-SGDJamie Hayes, Borja Balle, Saeed MahloujifarNeurIPS 2023 · 被引用 73 次
- PEARL: Data Synthesis via Private Embeddings and Adversarial Reconstruction LearningSeng Pei Liew, Tsubasa Takahashi, Michihiko UenoICLR 2022 · 被引用 32 次
- The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGDMilad Nasr, Thomas Steinke, Borja Balle, Christopher A. Choquette-Choo 等ICLR 2025
- Understanding Gradient Clipping in Private SGD: A Geometric PerspectiveXiangyi Chen, Zhiwei Steven Wu, Mingyi HongNeurIPS 2020 · 被引用 254 次
