SC2025Top-tier venue
Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
Pengfei Yu, Jingjing Gu, Hao Han, Dazhong Shen, Bao Wen, Yang Liu
Abstract
The exponential growth of Large Language Model (LLM) training demands in HPC systems has exposed critical reliability challenges, particularly from transient faults. Unlike resilience studies in conventional DNN inference, the massive parameter scale and iterative updates in LLM training trigger more complex failure patterns. To address these challenges, we introduce LLMFI, a new fault injection tool, and reveal six distinct failure behaviors through 300K+ fault injection experiments (exceeding 5K GPU node-hours). Our key insight is that, while most injected faults are eventually masked by the training iteration mechanism, a critical subset leads to catastrophic failures or performance degradation. Further, we propose LLMFT, a novel machine-learning-based fault tolerance framework that implements closed-loop error control via heuristic feature extraction, fault detector, and dual recovery mechanisms. Extensive evaluation demonstrates that LLMFT achieves an average of 97.61% F1-score in fault detection with only 0.01%–0.05% additional GPU memory overhead, effectively mitigating LLM training failures.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 66747f5f-30e9-4002-a8a9-5174aa534acbRelated papers
- ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model TrainingYuhang Liang, Xinyi Li, Jie Ren, Ang Li et al.PPoPP 2025 · 10 citations
- ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault ToleranceTong Xie, Jiawang Zhao, Zishen Wan, Zuodong Zhang et al.DAC 2025 · 4 citations
- Robust LLM Training Infrastructure at ByteDanceBorui Wan, Gaohong Liu, Zuquan Song, Jun Wang et al.SOSP 2025 · 1 citation
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun et al.NeurIPS 2025 · 1 citation
