A Unified Fast Gradient Clipping Framework for DP-SGD
Weiwei Kong, Andrés Muñoz Medina
Abstract
A well-known numerical bottleneck in the differentially-private stochastic gradient descent (DP-SGD) algorithm is the computation of the gradient norm for each example in a large input batch. When the loss function in DP-SGD consists of an intermediate linear operation, existing methods in the literature have proposed decompositions of gradients that are amenable to fast norm computations. In this paper, we present a framework that generalizes the above approach to arbitrary (possibly nonlinear) intermediate operations. Moreover, we show that for certain operations, such as fully-connected and embedding layer computations, further improvements to the runtime and storage costs of existing decompositions can be deduced using certain components of our framework. Finally, preliminary numerical experiments are given to demonstrate the substantial effects of the aforementioned improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c6aa3ac-c621-4775-8208-db17d412de86Cited by top-tier papers4
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway et al.ICML 2024 · 30 citations
- Per-example Gradients: a New Frontier for Understanding and Improving OptimizersVincent Roulet, Atish AgarwalaICML 2026 · 2 citations
- Differentially Private 2D Human Pose EstimationKaushik Bhargav Sivangi, Paul Henderson, Fani DeligianniCVPR 2026 · 1 citation
- Data Shapley in One Training RunJiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi JiaICLR 2025
Builds on8
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Reconstructing Training Data with Informed AdversariesBorja Balle, Giovanni Cherubin, Jamie HayesS&P 2022 · 214 citations
- Differentially Private Optimization on Large Model at Small CostZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisICML 2023 · 85 citations
Related papers
- Fast and Memory Efficient Differentially Private-SGD via JL ProjectionsZhiqi Bu, Sivakanth Gopi, Janardhan Kulkarni, Yin Tat Lee et al.NeurIPS 2021 · 49 citations
- Label Robust and Differentially Private Linear Regression: Computational and Statistical EfficiencyXiyang Liu, Prateek Jain, Weihao Kong, Sewoong Oh et al.NeurIPS 2023 · 10 citations
- DIFF2: Differential Private Optimization via Gradient Differences for Nonconvex Distributed LearningTomoya Murata, Taiji SuzukiICML 2023 · 11 citations
- Sparsity-Preserving Differentially Private Training of Large Embedding ModelsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar et al.NeurIPS 2023 · 9 citations
- Enabling Fast Differentially Private SGD via Just-in-Time Compilation and VectorizationPranav Subramani, Nicholas Vadivelu, Gautam KamathNeurIPS 2021 · 96 citations
