DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
Mogens Henrik From, Jacob Nielsen, Lukas Galke, Peter Schneider-Kamp
摘要
Training large neural network models requires extensive computational resources, often distributed across several nodes and accelerators. Recent findings suggest that it may be sufficient to only exchange the fast moving components of the gradients, while accumulating momentum locally (Decoupled Momentum, or DeMo). However, DeMo assumes that models fit on a single accelerator. We relax this assumption and introduce FlexDeMo, whereby nodes fully shard model parameters locally between different accelerators, while inter-node communication is reduced by synchronizing only fast-moving components instead of the full gradients -- resulting in a hybrid sharded data parallel training strategy. We further introduce a framework, denoted as DeToNATION, that generalizes DeMo, FlexDeMo, and other popular distributed training schemes such as DiLoCo -- introducing new variations of replication schemes and challenging choices made in DeMo. Our results across language and vision domains show that FlexDeMo attains similar validation loss as hybrid sharded data parallel training employing AdamW and full gradient synchronization, while being substantially faster. FlexDeMo is thus a promising distributed training scheme for the largest machine learning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- Efficient Sign-Based Optimization: Accelerating Convergence via Variance ReductionWei Jiang, Sifan Yang, Wenhao Yang, Lijun ZhangNeurIPS 2024 · 被引用 19 次
- DeMo: Decoupled Momentum OptimizationBowen Peng, Lizhang Chen, Baiyu Su, Jeffrey Quesnelle 等ICLR 2026 · 被引用 7 次
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk 等ACL 2024
相关 Paper
- MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local UpdatesAlex Iacob, Andrej Jovanovic, Mher Safaryan, Meghdad Kurmanji 等ICLR 2026 · 被引用 4 次
- DES-LOC: Desynced Low Communication Adaptive Optimizers for Foundation ModelsAlex Iacob, Lorenzo Sani, Mher Safaryan, Paris Giampouras 等ICLR 2026 · 被引用 2 次
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 被引用 9 次
- Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DMLJinkun Geng, Dan Li, Shuai WangICDE 2020 · 被引用 5 次
- Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCoChamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Hadi Mohaghegh Dolatabadi 等ICML 2026
