Sparse Checkpointing for Fast and Reliable MoE Training
Swapnil Gandhi, Christos Kozyrakis
摘要
As large language models scale, training them requires thousands of GPUs over extended durations-making frequent failures an inevitable reality. While checkpointing remains the primary fault-tolerance mechanism, existing methods fall short when applied to Mixture-of-Experts (MoE) models. Due to their substantially larger training state, MoE models exacerbate checkpointing overheads, often causing costly stalls or prolonged recovery that severely degrade training efficiency.
We present MoEvement, a distributed, in-memory checkpointing system tailored for MoE models. MoEvement is built on three key ideas: (1) sparse checkpointing, which incrementally snapshots subsets of experts across iterations to reduce overhead; (2) a sparse-to-dense checkpoint conversion mechanism that incrementally reconstructs consistent dense checkpoints from sparse snapshots; and (3) upstream logging of activations and gradients at pipeline-stage boundaries, enabling localized recovery without re-executing unaffected workers. Evaluations across diverse MoE models with up to 64 experts show that MoEvement reduces checkpointing overhead by up to 4× and recovery overhead by up to 31× compared to state-of-the-art approaches, sustaining ETTR ≥ 0.94 even under frequent failures (MTBF as low as 10 minutes) and delivering up to 8× overall training speedup, all without compromising synchronous training semantics. Overall, MoEvement offers a robust and scalable fault-tolerance solution for the next generation of sparsely activated models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
相关 Paper
- MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model TrainingWeilin Cai, Le Qin, Jiayi HuangASPLOS 2025 · 被引用 5 次
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi 等ASPLOS 2025 · 被引用 12 次
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsJiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang 等PPoPP 2022 · 被引用 97 次
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingAthinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie 等NSDI 2026 · 被引用 6 次
- Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE ParallelismChenwei Cui, Rockwell Jackson, Benjamin Joseph Herrera, Ana Tarano 等ICML 2026 · 被引用 1 次
