Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts
Max Ryabinin, Anton Gusev
Abstract
Many recent breakthroughs in deep learning were achieved by training increasingly larger models on massive datasets. However, training such models can be prohibitively expensive. For instance, the cluster used to train GPT-3 costs over $250 million. As a result, most researchers cannot afford to train state of the art models and contribute to their development. Hypothetically, a researcher could crowdsource the training of large neural networks with thousands of regular PCs provided by volunteers. The raw computing power of a hundred thousand $2500 desktops dwarfs that of a $250M server pod, but one cannot utilize that power efficiently with conventional distributed training methods. In this work, we propose Learning@home: a novel neural network training paradigm designed to handle large amounts of poorly connected participants. We analyze the performance, reliability, and architectural constraints of this paradigm and compare it against existing distributed training techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fae96d3-d158-4aeb-b141-407d214f6331Cited by top-tier papers18
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang et al.NeurIPS 2022 · 157 citations
- Model Ratatouille: Recycling Diverse Models for Out-of-Distribution GeneralizationAlexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord et al.ICML 2023 · 108 citations
- Distributed Deep Learning In Open CollaborationsMichael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier et al.NeurIPS 2021 · 89 citations
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 82 citations
- SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-EfficientMax Ryabinin, Tim Dettmers, Michael Diskin, Alexander BorzunovICML 2023 · 63 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- The HSIC Bottleneck: Deep Learning without Back-PropagationKurt Wan-Duo Ma, J. P. Lewis, W. Bastiaan KleijnAAAI 2020 · 180 citations
Related papers
- Varuna: scalable, low-cost training of massive deep learning modelsSanjith Athlur, Nitika Saran, Muthian Sivathanu, Ramachandran Ramjee et al.EuroSys 2022 · 81 citations
- Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingHaotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo et al.EMNLP 2025 · 2 citations
- Decentralized Diffusion ModelsDavid McAllister, Matthew Tancik, Jiaming Song, Angjoo KanazawaCVPR 2025
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Machine Learning on Volatile InstancesXiaoxi Zhang, Jianyu Wang, Gauri Joshi, Carlee Joe-WongINFOCOM 2020 · 17 citations
