Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data
Yasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu, Ethan Herron, Chinmay Hegde, Soumik Sarkar
Abstract
Decentralized learning enables a group of collaborative agents to learn models using a distributed dataset without the need for a central parameter server. Recently, decentralized learning algorithms have demonstrated state-of-the-art results on benchmark data sets, comparable with centralized algorithms. However, the key assumption to achieve competitive performance is that the data is independently and identically distributed (IID) among the agents which, in real-life applications, is often not applicable. Inspired by ideas from continual learning, we propose Cross-Gradient Aggregation (CGA), a novel decentralized learning algorithm where (i) each agent aggregates cross-gradient information, i.e., derivatives of its model with respect to its neighbors' datasets, and (ii) updates its model using a projected gradient based on quadratic programming (QP). We theoretically analyze the convergence characteristics of CGA and demonstrate its efficiency on non-IID data distributions sampled from the MNIST and CIFAR-10 datasets. Our empirical comparisons show superior learning performance of CGA over existing state-of-the-art decentralized learning algorithms, as well as maintaining the improved performance under information compression to reduce peer-to-peer communication overhead. The code is available here on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 284ea65e-6a34-40c4-9655-0cd881a875e2Cited by top-tier papers7
- Learning to Collaborate in Decentralized Learning of Personalized ModelsShuangtong Li, Tianyi Zhou, Xinmei Tian, Dacheng TaoCVPR 2022 · 41 citations
- Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous DataSai Aparna Aketi, Abolfazl Hashemi, Kaushik RoyNeurIPS 2023 · 23 citations
- Decentralized Dynamic Cooperation of Personalized Models for Federated Continual LearningDanni Yang, Zhikang Chen, Sen Cui, Mengyue Yang et al.NeurIPS 2025 · 2 citations
- DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning ModelsNastaran Saadati, Minh Pham, Nasla Saleem, Joshua R. Waite et al.CVPR 2024 · 2 citations
- GradMA: A Gradient-Memory-based Accelerated Federated Learning with Alleviated Catastrophic ForgettingKangyang Luo, Xiang Li, Yunshi Lan, Ming GaoCVPR 2023
Builds on6
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 672 citations
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
- Moniqua: Modulo Quantized Communication in Decentralized SGDYucheng Lu, Christopher De SaICML 2020 · 53 citations
Related papers
- Robust Distributed Gradient Aggregation Using Projections onto Gradient ManifoldsKwang In KimAAAI 2024
- Decentralized Learning with Multi-Headed DistillationAndrey Zhmoginov, Mark Sandler, Nolan Miller, Gus Kristiansen et al.CVPR 2023
- Robust Combination of Distributed Gradients Under Adversarial PerturbationsKwang In KimCVPR 2022 · 2 citations
- Decentralized Sporadic Federated Learning: A Unified Algorithmic Framework with Convergence GuaranteesShahryar Zehtabi, Dong-Jun Han, Rohit Parasnis, Seyyedali Hosseinalipour et al.ICLR 2025
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 95 citations
