A Theory to Instruct Differentially-Private Learning via Clipping Bias Reduction
Hanshen Xiao, Zihang Xiang, Di Wang, Srinivas Devadas
摘要
We study the bias introduced in Differentially-Private Stochastic Gradient Descent (DP-SGD) with clipped or normalized per-sample gradient. As one of the most popular but artificial operations to ensure bounded sensitivity, gradient clipping enables composite privacy analysis of many iterative optimization methods without additional assumptions on either learning models or input data. Despite its wide applicability, gradient clipping also presents theoretical challenges in systematically instructing improvement of privacy or utility. In general, without an assumption on globally-bounded gradient, classic convergence analyses do not apply to clipped gradient descent. Further, given limited understanding of the utility loss, many existing improvements to DP-SGD are heuristic, especially in the applications of private deep learning.In this paper, we provide meaningful theoretical analysis validated by thorough empirical results of DP-SGD. We point out that the bias caused by gradient clipping is underestimated in previous works. For generic non-convex optimization via DP-SGD, we show one key factor contributing to the bias is the sampling noise of stochastic gradient to be clipped. Accordingly, we use the developed theory to build a series of improvements for sampling noise reduction from various perspectives. From an optimization angle, we study variance reduction techniques and propose inner-outer momentum. At the learning model (neural network) level, we propose several tricks to enhance network internal normalization and BatchClipping to carefully clip the gradient of a batch of samples. For data preprocessing, we provide theoretical justification of recently proposed improvements via data normalization and (self-)augmentation.Putting these systematic improvements together, private deep learning via DP-SGD can be significantly strengthened in many tasks. For example, in computer vision applications, with an (ϵ = 8, δ = 10−5) DP guarantee, we successfully train ResNet20 on CIFAR10 and SVHN with test accuracy 76.0% and 90.1%, respectively; for natural language processing, with (ϵ = 4, δ = 10−5), we successfully train a recurrent neural network on IMDb data with test accuracy 77.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DP-Mix: Mixup-based Data Augmentation for Differentially Private LearningWenxuan Bao, Francesco Pittaluga, Vijay Kumar B. G, Vincent BindschaedlerNeurIPS 2023 · 被引用 19 次
- Delving into Differentially Private TransformerYoulong Ding, Xueyang Wu, Yining Meng, Yonggang Luo 等ICML 2024 · 被引用 11 次
- Geometry of Sensitivity: Twice Sampling and Hybrid Clipping in Differential Privacy with Optimal Gaussian Noise and Application to Deep LearningHanshen Xiao, Jun Wan, Srinivas DevadasCCS 2023 · 被引用 9 次
- Revisiting Differentially Private ReLU RegressionMeng Ding, Mingxi Lei, Liyang Zhu, Shaowei Wang 等NeurIPS 2024 · 被引用 7 次
- Revisiting Differentially Private Hyper-parameter TuningZihang Xiang, Tianhao Wang, Cheng-Long Wang, Di WangNDSS 2026 · 被引用 7 次
它引用的顶会 Paper18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 被引用 425 次
相关 Paper
- Enhancing DPSGD via Per-Sample Momentum and Low-Pass FilteringXincheng Xu, Thilina Ranbaduge, Qing Wang, Thierry Rakotoarivelo 等AAAI 2026
- Understanding Gradient Clipping in Private SGD: A Geometric PerspectiveXiangyi Chen, Zhiwei Steven Wu, Mingyi HongNeurIPS 2020 · 被引用 254 次
- Improved Convergence of Differential Private SGD with Gradient ClippingHuang Fang, Xiaoyun Li, Chenglin Fan, Ping LiICLR 2023
- Automatic Clipping: Differentially Private Deep Learning Made Easier and StrongerZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisNeurIPS 2023 · 被引用 140 次
- Beyond Uniform Lipschitz Condition in Differentially Private OptimizationRudrajit Das, Satyen Kale, Zheng Xu, Tong Zhang 等ICML 2023 · 被引用 24 次
