Beyond A Single AI Cluster: A Survey of Decentralized LLM Training
Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Bowen Li, Ying Shen, Zhi Wang
摘要
The emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations. Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities. As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches. We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements. We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA CommunicationMikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen 等SC 2025 · 被引用 6 次
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification SkippingHaotian Dong, Ye Li, Rongwei Lu, Chen Tang 等CVPR 2026 · 被引用 4 次
- DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE InferenceJiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song 等ICCV 2025 · 被引用 1 次
- Peak-Detector: Explainable Peak Detection via Instruction-Tuned Large Language Models in Physiological SignalJiahui Li, Yida Zhang, Zixuan Zeng, Jiayu Chen 等UbiComp 2026 · 被引用 1 次
- Di-PS: System-Algorithm Co-Design for Asynchronous and Heterogeneous Cross-cluster LLM Training at ScaleShengwei Li, Qiaoling Chen, Zhiquan Lai, Penglong Jiao 等NSDI 2026
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
相关 Paper
- vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingJehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim 等MICRO 2024 · 被引用 16 次
- Decentralized Diffusion ModelsDavid McAllister, Matthew Tancik, Jiaming Song, Angjoo KanazawaCVPR 2025
- A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with CrystalLLMShaoke Xi, ChonLam Lao, Boyi Jia, Jiaqi Gao 等SOSP 2026
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-ExpertsMax Ryabinin, Anton GusevNeurIPS 2020 · 被引用 71 次
