SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
Jin Lee, Zhonghao Chen, Xuhang He, Robert Underwood, Bogdan Nicolae, Franck Cappello, Xiaoyi Lu, Sheng Di, Zheng Zhang
Abstract
In large-scale LLM pretraining systems with GPUs, failures become the norm rather than the exception, and restart costs can dominate wall-clock training time. However, existing fault-tolerance mechanisms are largely unprepared for this restart-dominant regime. To address this challenge, we propose SPARe—Stacked Parallelism with Adaptive Reordering—a fault-tolerance framework that masks node failures during gradient synchronization by stacking redundant data shards across parallelism groups and adaptively reordering execution. SPARe achieves availability comparable to traditional replication while maintaining near-constant computation overhead of only , even under high redundancy where traidional replication would require linearly inflating overhead. We derive closed-form expressions for endurable failure count and computation overhead, validate them via SimGrid-based discrete-event simulation, and jointly optimize redundancy and checkpointing to minimize training time. At extreme scale with up to GPUs, SPARe reduces time-to-train by % compared to traditional replication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3bfdfc9-84f0-4866-992a-7ab5fb513309Builds on9
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- Calculon: a methodology and tool for high-level co-design of systems and large language modelsMikhail Isaev, Nic McDonald, Larry Dennison, Richard W. VuducSC 2023 · 42 citations
- GPU lifetimes on titan supercomputer: survival analysis and reliabilityGeorge Ostrouchov, Don Maxwell, Rizwan A. Ashraf, Christian Engelmann et al.SC 2020 · 38 citations
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello et al.HPDC 2024 · 34 citations
Related papers
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park et al.ICLR 2026
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang et al.OSDI 2026
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge et al.FAST 2026 · 3 citations
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
