Banded Square Root Matrix Factorization for Differentially Private Model Training
Nikita P. Kalinin, Christoph H. Lampert
Abstract
Current state-of-the-art methods for differentially private model training are based on matrix factorization techniques. However, these methods suffer from high computational overhead because they require numerically solving a demanding optimization problem to determine an approximately optimal factorization prior to the actual model training. In this work, we present a new matrix factorization approach, BSR, which overcomes this computational bottleneck. By exploiting properties of the standard matrix square root, BSR allows to efficiently handle also large-scale problems. For the key scenario of stochastic gradient descent with momentum and weight decay, we even derive analytical expressions for BSR that render the computational overhead negligible. We prove bounds on the approximation quality that hold both in the centralized and in the federated learning setting. Our numerical experiments demonstrate that models trained using BSR perform on par with the best existing methods, while completely avoiding their computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01148a71-55e6-4b80-a289-cbd870ce9bedCited by top-tier papers5
- Back to Square Roots: An Optimal Bound on the Matrix Factorization Error for Multi-Epoch Differentially Private SGDNikita Kalinin, Ryan McKenna, Jalaj Upadhyay, Christoph H. LampertICLR 2026 · 10 citations
- Continual Release Moment Estimation with Differential PrivacyNikita P. Kalinin, Jalaj Upadhyay, Christoph H. LampertNeurIPS 2025 · 5 citations
- Unified Privacy Guarantees for Decentralized Learning via Matrix FactorizationAurélien Bellet, Edwige Cyffers, Davide Frey, Romaric Gaudel et al.ICLR 2026 · 3 citations
- Scaling up the Banded Matrix Factorization Mechanism for Large Scale Differentially Private MLRyan McKennaICLR 2025
- Edit-Neighboring Data Streams and Privacy under Continual ObservationJoel Daniel Andersson, Anamay Chaturvedi, Monika Henzinger, Roodabeh SafaviCCS 2026
Builds on12
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Practical and Private (Deep) Learning Without Sampling or ShufflingPeter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar et al.ICML 2021 · 239 citations
- Hyperparameter Tuning with Renyi Differential PrivacyNicolas Papernot, Thomas SteinkeICLR 2022 · 157 citations
- Improved Differential Privacy for SGD via Optimal Private Linear Operators on Adaptive StreamsSergey Denisov, H. Brendan McMahan, John Rush, Adam D. Smith et al.NeurIPS 2022 · 96 citations
- (Amplified) Banded Matrix Factorization: A unified approach to private trainingChristopher A. Choquette-Choo, Arun Ganesh, Ryan McKenna, H. Brendan McMahan et al.NeurIPS 2023 · 67 citations
Related papers
- Gradient Descent with Linearly Correlated Noise: Theory and Applications to Differential PrivacyAnastasia Koloskova, Ryan McKenna, Zachary Charles, John Keith Rush et al.NeurIPS 2023 · 24 citations
- A Unified Fast Gradient Clipping Framework for DP-SGDWeiwei Kong, Andrés Muñoz MedinaNeurIPS 2023 · 9 citations
- Multi-Epoch Matrix Factorization Mechanisms for Private Machine LearningChristopher A. Choquette-Choo, Hugh Brendan McMahan, J. Keith Rush, Abhradeep Guha ThakurtaICML 2023 · 62 citations
- Sparsity-Preserving Differentially Private Training of Large Embedding ModelsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar et al.NeurIPS 2023 · 9 citations
- Federated Binary Matrix Factorization Using Proximal OptimizationSebastian Dalleiger, Jilles Vreeken, Michael KampAAAI 2025 · 1 citation
