Understanding Gradient Clipping in Private SGD: A Geometric Perspective
Xiangyi Chen, Zhiwei Steven Wu, Mingyi Hong
Abstract
Deep learning models are increasingly popular in many machine learning applications where the training data may contain sensitive information. To provide formal and rigorous privacy guarantee, many learning systems now incorporate differential privacy by training their models with (differentially) private SGD. A key step in each private SGD update is gradient clipping that shrinks the gradient of an individual example whenever its 2 norm exceeds some threshold. We first demonstrate how gradient clipping can prevent SGD from converging to a stationary point. We then provide a theoretical analysis that fully quantifies the clipping bias on convergence with a disparity measure between the gradient distribution and a geometrically symmetric distribution. Our empirical evaluation further suggests that the gradient distributions along the trajectory of private SGD indeed exhibit symmetric structure that favors convergence. Together, our results provide an explanation why private SGD with gradient clipping remains effective in practice despite its potential clipping bias. Finally, we develop a new perturbation-based technique that can provably correct the clipping bias even for instances with highly asymmetric gradient distributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers62
- Automatic Clipping: Differentially Private Deep Learning Made Easier and StrongerZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisNeurIPS 2023 · 140 citations
- Understanding Clipping for Federated Learning: Convergence and Client-Level Differential PrivacyXinwei Zhang, Xiangyi Chen, Mingyi Hong, Steven Wu et al.ICML 2022 · 134 citations
- Do not Let Privacy Overbill Utility: Gradient Embedding Perturbation for Private LearningDa Yu, Huishuai Zhang, Wei Chen, Tie-Yan LiuICLR 2021 · 133 citations
- Revisiting Gradient Clipping: Stochastic bias and tight convergence guaranteesAnastasia Koloskova, Hadrien Hendrikx, Sebastian U. StichICML 2023 · 106 citations
- What Do We Mean by Generalization in Federated Learning?Honglin Yuan, Warren Richard Morningstar, Lin Ning, Karan SinghalICLR 2022 · 98 citations
Builds on3
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 425 citations
Related papers
- GeoClip: Geometry-Aware Clipping for Differentially Private SGDAtefeh Gilani, Naima Tasnim, Lalitha Sankar, Oliver KosutNeurIPS 2025 · 5 citations
- Differentially Private SGD Without Clipping Bias: An Error-Feedback ApproachXinwei Zhang, Zhiqi Bu, Steven Wu, Mingyi HongICLR 2024 · 15 citations
- A Theory to Instruct Differentially-Private Learning via Clipping Bias ReductionHanshen Xiao, Zihang Xiang, Di Wang, Srinivas DevadasS&P 2023
- Improved Convergence of Differential Private SGD with Gradient ClippingHuang Fang, Xiaoyun Li, Chenglin Fan, Ping LiICLR 2023
- Adaptive Sigmoid Clipping for Balancing the Direction-Magnitude Mismatch Trade-off in Differentially Private LearningFaeze Moradi Kalarde, Ali Bereyhi, Ben Liang, Min DongNeurIPS 2025 · 1 citation
