SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular Redundancy
Rihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park, Jae W. Lee
摘要
Dual Modular Redundancy (DMR) is a highly effective mechanism for detecting silent data corruption (SDC)—a critical reliability concern in large language model (LLM) training—by executing each operation twice. However, its high computation overhead has prevented practical deployment at scale. In this paper, we present SpareTrain, an LLM training system that achieves complete DMR with minimal overhead by repurposing the activation checkpointing mechanism and exploiting idle GPU time. Evaluations on up to 32 H200 GPUs show that SpareTrain improves throughput by 12–35% over naive DMR, corresponding to only 3–14% overhead compared to unprotected training, while maintaining full DMR error detection capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis 等ASPLOS 2023 · 被引用 64 次
- Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsYi He, Mike Hutton, Steven Chan, Robert De Gruijl 等ISCA 2023 · 被引用 52 次
- Calculon: a methodology and tool for high-level co-design of systems and large language modelsMikhail Isaev, Nic McDonald, Larry Dennison, Richard W. VuducSC 2023 · 被引用 42 次
相关 Paper
- SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUsJin Lee, Zhonghao Chen, Xuhang He, Robert Underwood 等ICML 2026 · 被引用 1 次
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen 等OSDI 2026
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge 等FAST 2026 · 被引用 3 次
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu 等DAC 2025 · 被引用 3 次
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
