Lune

ICLR2025顶会

Adaptive Gradient Clipping for Robust Federated Learning

Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Ahmed Jellouli, Geovani Rizk, John Stephan

出版方
2025年份
2顶会引用

摘要

Robust federated learning aims to maintain reliable performance despite the presence of adversarial or misbehaving workers. While state-of-the-art (SOTA) robust distributed gradient descent (Robust-DGD) methods were proven theoretically optimal, their empirical success has often relied on pre-aggregation gradient clipping. However, existing static clipping strategies yield inconsistent results: enhancing robustness against some attacks while being ineffective or even detrimental against others. To address this limitation, we propose a principled adaptive clipping strategy, Adaptive Robust Clipping (ARC), which dynamically adjusts clipping thresholds based on the input gradients. We prove that ARC not only preserves the theoretical robustness guarantees of SOTA Robust-DGD methods but also provably improves asymptotic convergence when the model is well-initialized. Extensive experiments on benchmark image classification tasks confirm these theoretical insights, demonstrating that ARC significantly enhances robustness, particularly in highly heterogeneous and adversarial settings.

  • Authors are listed in alphabetical order.

Published as a conference paper at ICLR 2025 settings and adversarial regimes. Our results demonstrate that ARC significantly enhances the performance of state-of-the-art Robust-DGD methods, particularly in scenarios with high data heterogeneity (Figure 1a) and a large number of adversarial workers (Figure 4b).

(3) Improved learning guarantee. We demonstrate that ARC possesses an additional property that is not satisfied by classical robust aggregation methods. Specifically, ARC constrains the norm of an adversarial gradient by that of an honest (non-adversarial) gradient. Leveraging this property, we show that ARC circumvents the lower bound established under data heterogeneity in Allouah et al. (2023b), provided the honest gradients are bounded at model initialization. An empirical validation of this insight is shown in Figure 1b. Such model initialization is often satisfiable in practice (Glorot & Bengio, 2010), highlighting the practical relevance of ARC. When the model is arbitrarily initialized, ARC recovers the original convergence guarantee of Robust-DGD in the worst case.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖