SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular Redundancy
Rihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park, Jae W. Lee
Abstract
Dual Modular Redundancy (DMR) is a highly effective mechanism for detecting silent data corruption (SDC)—a critical reliability concern in large language model (LLM) training—by executing each operation twice. However, its high computation overhead has prevented practical deployment at scale. In this paper, we present SpareTrain, an LLM training system that achieves complete DMR with minimal overhead by repurposing the activation checkpointing mechanism and exploiting idle GPU time. Evaluations on up to 32 H200 GPUs show that SpareTrain improves throughput by 12–35% over naive DMR, corresponding to only 3–14% overhead compared to unprotected training, while maintaining full DMR error detection capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 917ebb8c-5feb-4359-8418-e9ee10db709fBuilds on13
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet et al.ASPLOS 2022 · 68 citations
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis et al.ASPLOS 2023 · 64 citations
- Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsYi He, Mike Hutton, Steven Chan, Robert De Gruijl et al.ISCA 2023 · 52 citations
- Calculon: a methodology and tool for high-level co-design of systems and large language modelsMikhail Isaev, Nic McDonald, Larry Dennison, Richard W. VuducSC 2023 · 42 citations
Related papers
- SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUsJin Lee, Zhonghao Chen, Xuhang He, Robert Underwood et al.ICML 2026 · 1 citation
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen et al.OSDI 2026
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge et al.FAST 2026 · 3 citations
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu et al.DAC 2025 · 3 citations
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
