Adaptive Gradient Clipping for Robust Federated Learning
Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Ahmed Jellouli, Geovani Rizk, John Stephan
摘要
Robust federated learning aims to maintain reliable performance despite the presence of adversarial or misbehaving workers. While state-of-the-art (SOTA) robust distributed gradient descent (Robust-DGD) methods were proven theoretically optimal, their empirical success has often relied on pre-aggregation gradient clipping. However, existing static clipping strategies yield inconsistent results: enhancing robustness against some attacks while being ineffective or even detrimental against others. To address this limitation, we propose a principled adaptive clipping strategy, Adaptive Robust Clipping (ARC), which dynamically adjusts clipping thresholds based on the input gradients. We prove that ARC not only preserves the theoretical robustness guarantees of SOTA Robust-DGD methods but also provably improves asymptotic convergence when the model is well-initialized. Extensive experiments on benchmark image classification tasks confirm these theoretical insights, demonstrating that ARC significantly enhances robustness, particularly in highly heterogeneous and adversarial settings.
- Authors are listed in alphabetical order.
Published as a conference paper at ICLR 2025 settings and adversarial regimes. Our results demonstrate that ARC significantly enhances the performance of state-of-the-art Robust-DGD methods, particularly in scenarios with high data heterogeneity (Figure 1a) and a large number of adversarial workers (Figure 4b).
(3) Improved learning guarantee. We demonstrate that ARC possesses an additional property that is not satisfied by classical robust aggregation methods. Specifically, ARC constrains the norm of an adversarial gradient by that of an honest (non-adversarial) gradient. Leveraging this property, we show that ARC circumvents the lower bound established under data heterogeneity in Allouah et al. (2023b), provided the honest gradients are bounded at model initialization. An empirical validation of this insight is shown in Figure 1b. Such model initialization is often satisfiable in practice (Glorot & Bengio, 2010), highlighting the practical relevance of ARC. When the model is arbitrarily initialized, ARC recovers the original convergence guarantee of Robust-DGD in the worst case.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Tight Stability Bounds for Robust Distributed Learning: Byzantine Failures Hurt Generalization More than Data PoisoningThomas Boudou, Batiste Le Bars, Nirupam Gupta, Aurélien BelletICML 2026 · 被引用 3 次
- Unified Breakdown Analysis for Byzantine Robust GossipRenaud Gaucher, Aymeric Dieuleveut, Hadrien HendrikxICML 2025
它引用的顶会 Paper19
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 被引用 247 次
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 被引用 192 次
相关 Paper
- Byzantine-Robust Learning on Heterogeneous Data via Gradient SplittingYuchen Liu, Chen Chen, Lingjuan Lyu, Fangzhao Wu 等ICML 2023 · 被引用 27 次
- Robust Federated Learning: The Case of Affine Distribution ShiftsAmirhossein Reisizadeh, Farzan Farnia, Ramtin Pedarsani, Ali JadbabaieNeurIPS 2020 · 被引用 196 次
- Noise-Aware Algorithm for Heterogeneous Differentially Private Federated LearningSaber Malekmohammadi, Yaoliang Yu, Yang CaoICML 2024 · 被引用 10 次
- Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker AssumptionsWeixin An, Yuanyuan Liu, Fanhua Shang, Han Yu 等NeurIPS 2025
- FedGPS: Statistical Rectification Against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Chenxin Li, Yiu-ming Cheung 等NeurIPS 2025 · 被引用 7 次
