The Non-IID Data Quagmire of Decentralized Machine Learning
Kevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. Gibbons
Abstract
Many large-scale machine learning (ML) applications need to perform decentralized learning over datasets generated at different devices and locations. Such datasets pose a significant challenge to decentralized learning because their different contexts result in significant data distribution skew across devices/locations. In this paper, we take a step toward better understanding this challenge by presenting a detailed experimental study of decentralized DNN training on a common type of data skew: skewed distribution of data labels across devices/locations. Our study shows that: (i) skewed data labels are a fundamental and pervasive problem for decentralized learning, causing significant accuracy loss across many ML applications, DNN models, training datasets, and decentralized learning algorithms; (ii) the problem is particularly challenging for DNN models with batch normalization; and (iii) the degree of data skew is a key determinant of the difficulty of the problem. Based on these findings, we present SkewScout, a system-level approach that adapts the communication frequency of decentralized learning algorithms to the (skew-induced) accuracy loss between data partitions. We also show that group normalization can recover much of the accuracy loss of batch normalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8ea5977-9e87-4954-97dd-442ec5e98c3aCited by top-tier papers102
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp et al.ICLR 2021 · 1,166 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 605 citations
Builds on3
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Fair Resource Allocation in Federated LearningTian Li, Maziar Sanjabi, Ahmad Beirami, Virginia SmithICLR 2020 · 971 citations
Related papers
- Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous DataSai Aparna Aketi, Abolfazl Hashemi, Kaushik RoyNeurIPS 2023 · 23 citations
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi et al.ICML 2021 · 100 citations
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 118 citations
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang et al.ICCV 2021 · 73 citations
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li et al.NeurIPS 2021 · 61 citations
