Taming Fat-Tailed ("Heavier-Tailed" with Potentially Infinite Variance) Noise in Federated Learning
Haibo Yang, Peiwen Qiu, Jia Liu
摘要
In recent years, federated learning (FL) has emerged as an important distributed machine learning paradigm to collaboratively learn a global model with multiple clients, while keeping data local and private. However, a key assumption in most existing works on FL algorithms' convergence analysis is that the noise in stochastic first-order information has a finite variance. Although this assumption covers all light-tailed (i.e., sub-exponential) and some heavy-tailed noise distributions (e.g., log-normal, Weibull, and some Pareto distributions), it fails for many fat-tailed noise distributions (i.e., "heavier-tailed" with potentially infinite variance) that have been empirically observed in the FL literature. To date, it remains unclear whether one can design convergent algorithms for FL systems that experience fat-tailed noise. This motivates us to fill this gap in this paper by proposing an algorithmic framework called FAT-Clipping (federated averaging with two-sided learning rates and clipping), which contains two variants: FAT-Clipping per-round (FAT-Clipping-PR) and FAT-Clipping per-iteration (FAT-Clipping-PI). Specifically, for the largest tail-index α ∈ (1, 2] such that the fat-tailed noise in FL still has a bounded α-moment, we show that both variants achieve O((mT ) 2-α α ) and O((mT ) 1-α 3α-2 ) convergence rates in the strongly-convex and general non-convex settings, respectively, where m and T are the numbers of clients and communication rounds. Moreover, with more clipping operations compared to FAT-Clipping-PR, FAT-Clipping-PI further enjoys a linear speedup effect with respect to the number of local updates at each client and being lower-bound-matching (i.e., order-optimal). Collectively, our results advance the understanding of designing efficient algorithms for FL systems that exhibit fat-tailed first-order oracle information. 2 Related work In this section, we will provide a quick overview on three related topics in the literature: i) federated learning, ii) heavy-tailed noise in learning, and iii) the clipping techniques, thus putting our work into comparative perspective to highlight our novelty and differences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Federated Multi-Objective LearningHaibo Yang, Zhuqing Liu, Jia Liu, Chaosheng Dong 等NeurIPS 2023 · 被引用 28 次
- An improved analysis of per-sample and per-update clipping in federated learningBo Li, Xiaowen Jiang, Mikkel N. Schmidt, Tommy Sonne Alstrøm 等ICLR 2024 · 被引用 9 次
- Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal ConvergenceShuhua Yu, Dusan Jakovetic, Soummya KarICLR 2026 · 被引用 7 次
- Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker AssumptionsWeixin An, Yuanyuan Liu, Fanhua Shang, Han Yu 等NeurIPS 2025
- Efficient Distributed Optimization under Heavy-Tailed NoiseSu Hyeong Lee, Manzil Zaheer, Tian LiICML 2025
它引用的顶会 Paper24
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi 等NeurIPS 2020 · 被引用 2,231 次
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
相关 Paper
- Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient ClippingZijian Liu, Zhengyuan ZhouICLR 2025
- Understanding Clipping for Federated Learning: Convergence and Client-Level Differential PrivacyXinwei Zhang, Xiangyi Chen, Mingyi Hong, Steven Wu 等ICML 2022 · 被引用 134 次
- High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed NoiseEduard Gorbunov, Abdurakhmon Sadiev, Marina Danilova, Samuel Horváth 等ICML 2024 · 被引用 27 次
- Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGradZijian LiuICML 2026 · 被引用 3 次
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 被引用 5 次
