Efficient DP-SGD for LLMs with Randomized Clipping
Enayat Ullah, Sai Aparna Aketi, Devansh Gupta, Huanyu Zhang, Meisam Razaviyayn
Abstract
Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for formal privacy guarantees, provides a principled framework for training LLMs with provable privacy protection. However, state-of-the-art DP training implementations rely on fast gradient clipping techniques with memory overhead O(B minT 2 , d 2 ), where B is the batch size, T , the sequence length, and d, the model width. This becomes prohibitive as both model size and context length grow. We propose DP-SGD-RC, a novel variant of DP-SGD with randomized clipping that reduces memory and compute complexity. DP-SGD-RC leverages stochastic trace estimation methods, specifically Hutchinson's estimator [Hutchinson, 1989] and its improved variant, Hutch ++ [Meyer et al., 2021], to reduce the memory footprint of per-sample gradient norm estimation. We provide a tight privacy analysis showing that DP-SGD-RC achieves noise multipliers competitive with deterministic clipping. Experiments fine-tuning Llama 3.2 1B on long-context benchmarks spanning classification, question answering, and summarization tasks demonstrate that DP-SGD-RC matches baseline utility while significantly reducing memory and compute.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 425 citations
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 259 citations
Related papers
- Private Training Large-scale Models with Efficient DP-SGDLiangyu Wang, Junxiao Wang, Jie Ren, Zihang Xiang et al.NeurIPS 2025 · 8 citations
- Differentially Private Optimization on Large Model at Small CostZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisICML 2023 · 85 citations
- DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise ReductionXinwei Zhang, Zhiqi Bu, Borja Balle, Mingyi Hong et al.ICLR 2025
- Improved Convergence of Differential Private SGD with Gradient ClippingHuang Fang, Xiaoyun Li, Chenglin Fan, Ping LiICLR 2023
- Prism: Private Relational Data Synthesis with Language ModelsGuohui Guan, Chang GeSIGMOD 2026
