DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
Jin Liu, Ning Xi, Yinbin Miao, Junkang Liu
Abstract
Balancing convergence efficiency and robustness under Differential Privacy (DP) is a central challenge in Federated Learning (FL). While AdamW accelerates training and fine-tuning in large-scale models, we find that directly applying it to Differentially Private FL (DPFL) suffers from three major issues: (i) data heterogeneity and privacy noise jointly amplify the variance of second-moment estimator, (ii) DP perturbations bias the second-moment estimator, and (iii) DP amplify AdamW’s sensitivity to local overfitting, worsening client drift. We propose DP-FedAdamW, the first AdamW-based optimizer for DPFL. It restores AdamW under DP by stabilizing second-moment variance, removing DP-induced bias, and aligning local updates to the global descent to curb client drift.Theoretically, we establish an unbiased second-moment estimator and prove a linearly accelerated convergence rate without any heterogeneity assumption, while providing tighter -DP guarantees.Our empirical results demonstrate the effectiveness of DP-FedAdamW across language and vision Transformers and ResNet-18. On Tiny-ImageNet (Swin-Base, ), DP-FedAdamW outperforms the state-of-the-art (SOTA) by 5.83%. Thecode is available in Appendix.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
Related papers
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian et al.AAAI 2026 · 12 citations
- Rethinking LoRA for Privacy-Preserving Federated Learning in Large ModelsJin Liu, Yinbin Miao, Ning Xi, Junkang LiuICLR 2026 · 9 citations
- FIBER: A Differentially Private Optimizer with Filter-Aware Innovation Bias CorrectionMINH DUC DO, Thao Do, Minh Hoang, Anh Le Duc Tran et al.ICML 2026 · 1 citation
- Efficient and Differentially Private Federated LLM Fine-Tuning on Heterogeneous ClientsNan Yan, Yuqing Li, Xiong Wang, Jing Chen et al.KDD 2026
- Make Landscape Flatter in Differentially Private Federated LearningYifan Shi, Yingqi Liu, Kang Wei, Li Shen et al.CVPR 2023
