Consistency of Local and Global Flatness for Federated Learning
Junkang Liu, Fanhua Shang, Yuxuan Tian, Hongying Liu, Yuanyuan Liu
Abstract
In federated learning (FL), multi-step local updates and data heterogeneity usually lead to sharper global minima, which degrades the performance of the global model. Popular FL algorithms integrate sharpness-aware minimization (SAM) into local training to address this issue. However, in the high data heterogeneity setting, the flatness in local training does not imply the flatness of the global model. Therefore, minimizing the sharpness of the local loss surfaces on the client data does not enable the effectiveness of SAM in FL to improve the generalization ability of the global model. We define the flatness distance to explain this phenomenon. By rethinking the SAM in FL and theoretically analyzing the flatness distance, we propose a novel FedNSAM algorithm that accelerates the SAM algorithm by introducing global Nesterov momentum into the local update to harmonize the consistency of global and local flatness. FedNSAM uses the global Nesterov momentum as the direction of local estimation of client global perturbations and extrapolation. Theoretically, we prove a tighter convergence bound than FedSAM by Nesterov extrapolation. Empirically, we conduct comprehensive experiments on CNN and Transformer models to verify the superior performance and efficiency of FedNSAM. The code is available at https://github.com/junkangLiu0/FedNSAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af012fd9-45a7-430a-8cfe-3e60bc978a90Cited by top-tier papers12
- FedALT: Federated Fine-Tuning Through Adaptive Local Training with Rest-of-World LoRAJieming Bian, Lei Wang, Letian Zhang, Jie XuAAAI 2026 · 12 citations
- Rethinking LoRA for Privacy-Preserving Federated Learning in Large ModelsJin Liu, Yinbin Miao, Ning Xi, Junkang LiuICLR 2026 · 9 citations
- CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud TrackingSifan Zhou, Yichao Cao, Jiahao Nie, Yuqian Fu et al.AAAI 2026 · 9 citations
- FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated LearningJunkang Liu, Fanhua Shang, Yuanyuan Liu, Hongying Liu et al.ACM MM 2024 · 6 citations
- HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu et al.ACM MM 2025 · 5 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu et al.NeurIPS 2025 · 228 citations
Related papers
- Generalized Federated Learning via Sharpness Aware MinimizationZhe Qu, Xingyu Li, Rui Duan, Yao Liu et al.ICML 2022 · 219 citations
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu et al.ICML 2024 · 35 citations
- One Arrow, Two Hawks: Sharpness-aware Minimization for Federated Learning via Global Model TrajectoryYuhang Li, Tong Liu, Yangguang Cui, Ming Hu et al.ICML 2025
- FedAdamom: Adaptive Momentum for Improved Generalization in Federated OptimizationWenjie Hou, Tianxiang Chen, Feng Wang, Tiantong Wu et al.CVPR 2026
- Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight AveragingJunkang Liu, Yuanyuan Liu, Fanhua Shang, Hongying Liu et al.ICML 2025
