Distributed Deep Learning In Open Collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Quentin Lhoest, Anton Sinitsin, Dmitry Popov, Dmitry V. Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, Denis Mazur
摘要
Modern deep learning applications require increasingly more compute to train state-of-the-art models. To address this demand, large corporations and institutions use dedicated High-Performance Computing clusters, whose construction and maintenance are both environmentally costly and well beyond the budget of most organizations. As a result, some research directions become the exclusive domain of a few large industrial and even fewer academic actors. To alleviate this disparity, smaller groups may pool their computational resources and run collaborative experiments that benefit all participants. This paradigm, known as grid- or volunteer computing, has seen successful applications in numerous scientific areas. However, using this approach for machine learning is difficult due to high latency, asymmetric bandwidth, and several challenges unique to volunteer computing. In this work, we carefully analyze these constraints and propose a novel algorithmic framework designed specifically for collaborative training. We demonstrate the effectiveness of our approach for SwAV and ALBERT pretraining in realistic conditions and achieve performance comparable to traditional setups at a fraction of the cost. Finally, we provide a detailed report of successful collaborative language model pretraining with 40 participants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-EfficientMax Ryabinin, Tim Dettmers, Michael Diskin, Alexander BorzunovICML 2023 · 被引用 63 次
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen 等ICML 2023 · 被引用 60 次
- HexGen: Generative Inference of Large Language Model over Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Xiaozhe Yao, Yang Zhou 等ICML 2024 · 被引用 46 次
- Distributed Methods with Compressed Communication for Solving Variational Inequalities, with Theoretical GuaranteesAleksandr Beznosikov, Peter Richtárik, Michael Diskin, Max Ryabinin 等NeurIPS 2022 · 被引用 25 次
- Secure Distributed Training at ScaleEduard Gorbunov, Alexander Borzunov, Michael Diskin, Max RyabininICML 2022 · 被引用 18 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone 等CCS 2017 · 被引用 3,936 次
相关 Paper
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-ExpertsMax Ryabinin, Anton GusevNeurIPS 2020 · 被引用 71 次
- Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingHaotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo 等EMNLP 2025 · 被引用 2 次
- Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile DevicesYuhao Chen, Yuxuan Yan, Shuowei Ge, Yuyang Qin 等MobiCom 2025 · 被引用 4 次
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu 等ICML 2023 · 被引用 7 次
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable DevicesMax Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, Gennady PekhimenkoNeurIPS 2021 · 被引用 59 次
