Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC Systems
Pengfei Yu, Jingjing Gu, Hao Han, Dazhong Shen, Bao Wen, Yang Liu
摘要
The exponential growth of Large Language Model (LLM) training demands in HPC systems has exposed critical reliability challenges, particularly from transient faults. Unlike resilience studies in conventional DNN inference, the massive parameter scale and iterative updates in LLM training trigger more complex failure patterns. To address these challenges, we introduce LLMFI, a new fault injection tool, and reveal six distinct failure behaviors through 300K+ fault injection experiments (exceeding 5K GPU node-hours). Our key insight is that, while most injected faults are eventually masked by the training iteration mechanism, a critical subset leads to catastrophic failures or performance degradation. Further, we propose LLMFT, a novel machine-learning-based fault tolerance framework that implements closed-loop error control via heuristic feature extraction, fault detector, and dual recovery mechanisms. Extensive evaluation demonstrates that LLMFT achieves an average of 97.61% F1-score in fault detection with only 0.01%–0.05% additional GPU memory overhead, effectively mitigating LLM training failures.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model TrainingYuhang Liang, Xinyi Li, Jie Ren, Ang Li 等PPoPP 2025 · 被引用 10 次
- ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault ToleranceTong Xie, Jiawang Zhao, Zishen Wan, Zuodong Zhang 等DAC 2025 · 被引用 4 次
- Robust LLM Training Infrastructure at ByteDanceBorui Wan, Gaohong Liu, Zuquan Song, Jun Wang 等SOSP 2025 · 被引用 1 次
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun 等NeurIPS 2025 · 被引用 1 次
