Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]
Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang, Jingxian Wang, Wentao Wu, Qizhen Zhang
摘要
Federated learning (FL) has emerged as a popular paradigm for distributed machine learning over decentralized data. A typical FL training task involves a fleet of client devices with private data and a centralized server for aggregating the global model. Data generated by FL clients, e.g., smart phones, vehicles, and cameras, is prone to noise. While the impact of data noise on centralized learning (CL) is well understood, to our best knowledge there is a lack of a systematic study from this point of view for FL. In this paper, we fill this gap by presenting an empirical investigation to provide a deeper understanding regarding the impact of data noise on FL. Our study is enabled by DataNoiseGenerator, an open-source and extensible toolkit that we developed for the injection of controlled data noise across five diverse data modalities: image, video, audio, text, and tabular data. We then carry out extensive experiments based on the noisy data generated by DataNoiseGenerator, and our experimental evaluation results reveal that FL is significantly more vulnerable to data noise compared to CL, in terms of the quality of the trained ML models. This gap between FL and CL widens as the intensity of data noise and the proportion of noisy FL clients increase. We further present a detailed analysis to diagnose the root cause of this increased sensitivity of FL to data noise. Our analysis finds that the aggregation performed by the FL server can amplify divergent updates from FL clients trained on noisy data, thereby hindering global model convergence. We conclude that data quality issues are a fundamental challenge for deploying robust FL systems and demand novel decentralized data cleaning mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Robust Federated Learning with Noisy and Heterogeneous ClientsXiuwen Fang, Mang YeCVPR 2022 · 被引用 169 次
- Robust Heterogeneous Federated Learning under Data CorruptionXiuwen Fang, Mang Ye, Xiyuan YangICCV 2023 · 被引用 44 次
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu 等VLDB 2023 · 被引用 24 次
相关 Paper
- Dealing with Noisy Data in Federated Learning: An Incentive Mechanism with Flexible PricingHengzhi Wang, Haoran Chen, Minghe Ma, Laizhong CuiWWW 2025 · 被引用 3 次
- FedClean: A General Robust Label Noise Correction for Federated LearningXiaoqian Jiang, Jing ZhangICML 2025
- martFL: Enabling Utility-Driven Data Marketplace with a Robust and Verifiable Federated Learning ArchitectureQi Li, Zhuotao Liu, Qi Li, Ke XuCCS 2023 · 被引用 19 次
- SAFER-FL: Adversarially Robust Federated Learning in IoT Networks via Latent Space Auditing and Verifiable ContributionsNaveen Kumar Kummari, Mohsen GuizaniINFOCOM 2026
- FedCorr: Multi-Stage Federated Learning for Label Noise CorrectionJingyi Xu, Zihan Chen, Tony Q. S. Quek, Kai Fong Ernest ChongCVPR 2022 · 被引用 101 次
