Federated Learning on Non-IID Data Silos: An Experimental Study
Qinbin Li, Yiqun Diao, Quan Chen, Bingsheng He
摘要
Due to the increasing privacy concerns and data regulations, training data have been increasingly fragmented, forming distributed databases of multiple “data silos” (e.g., within different organizations and countries). To develop effective machine learning services, there is a must to exploit data from such distributed databases without exchanging the raw data. Recently, federated learning (FL) has been a solution with growing interests, which enables multiple parties to collaboratively train a machine learning model without exchanging their local data. A key and common challenge on distributed databases is the heterogeneity of the data distribution among the parties. The data of different parties are usually non-independently and identically distributed (i.e., non-IID). There have been many FL algorithms to address the learning effectiveness under non-IID data settings. However, there lacks an experimental study on systematically understanding their advantages and disadvantages, as previous studies have very rigid data partitioning strategies among parties, which are hardly representative and thorough. In this paper, to help researchers better understand and study the non-IID data setting in federated learning, we propose comprehensive data partitioning strategies to cover the typical non-IID data cases. Moreover, we conduct extensive experiments to evaluate state-of-the-art FL algorithms. We find that non-IID does bring significant challenges in learning accuracy of FL algorithms, and none of the existing state-of-the-art FL algorithms outperforms others in all cases. Our experiments provide insights for future studies of addressing the challenges in “data silos”.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper180
- No Fear of Heterogeneity: Classifier Calibration for Federated Learning with Non-IID DataMi Luo, Fei Chen, Dapeng Hu, Yifan Zhang 等NeurIPS 2021 · 被引用 510 次
- Stochastic Controlled Averaging for Federated Learning with Communication CompressionXinmeng Huang, Ping Li, Xiaoyun LiICLR 2024 · 被引用 288 次
- Learn from Others and Be Yourself in Heterogeneous Federated LearningWenke Huang, Mang Ye, Bo DuCVPR 2022 · 被引用 254 次
- Preservation of the Global Knowledge by Not-True Distillation in Federated LearningGihun Lee, Minchan Jeong, Yongjin Shin, Sangmin Bae 等NeurIPS 2022 · 被引用 235 次
- NOTE: Robust Continual Test-time Adaptation Against Temporal CorrelationTaesik Gong, Jongheon Jeong, Taewon Kim, Yewon Kim 等NeurIPS 2022 · 被引用 227 次
它引用的顶会 Paper24
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi 等NeurIPS 2020 · 被引用 2,231 次
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 被引用 1,615 次
相关 Paper
- DistFL: Distribution-aware Federated Learning for Mobile ScenariosBingyan Liu, Yifeng Cai, Ziqi Zhang, Yuanchun Li 等UbiComp 2022 · 被引用 19 次
- Distribution-Regularized Federated Learning on Non-IID DataYansheng Wang, Yongxin Tong, Zimu Zhou, Ruisheng Zhang 等ICDE 2023 · 被引用 31 次
- Adversarial Collaborative Learning on Non-IID FeaturesQinbin Li, Bingsheng He, Dawn SongICML 2023 · 被引用 22 次
- Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective AlignmentLin Zhang, Yong Luo, Yan Bai, Bo Du 等ICCV 2021 · 被引用 98 次
- FedCross: Towards Accurate Federated Learning via Multi-Model Cross-AggregationMing Hu, Peiheng Zhou, Zhihao Yue, Zhiwei Ling 等ICDE 2024 · 被引用 32 次
