ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
Yuhang Liang, Xinyi Li, Jie Ren, Ang Li, Bo Fang, Jieyang Chen
Abstract
Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, AT-TNChecker reduces recovery overhead by up to 49×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f9c22bd-3354-4823-a9bf-7950426bec0bCited by top-tier papers4
- FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant AttentionHuangliang Dai, Shixun Wu, Jiajun Huang, Zizhe Jian et al.SC 2025 · 12 citations
- Demystifying the Resilience of Large Language Model Inference: An End-to-End PerspectiveYu Sun, Zachary Coalson, Shiyang Chen, Hang Liu et al.SC 2025 · 9 citations
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park et al.ICLR 2026
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen et al.OSDI 2026
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsYi He, Mike Hutton, Steven Chan, Robert De Gruijl et al.ISCA 2023 · 52 citations
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng et al.NSDI 2025 · 46 citations
- Improving Energy Saving of One-Sided Matrix Decompositions on CPU-GPU Heterogeneous SystemsJieyang Chen, Xin Liang, Kai Zhao, Hadi Zamani Sabzi et al.PPoPP 2023 · 4 citations
Related papers
- Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC SystemsPengfei Yu, Jingjing Gu, Hao Han, Dazhong Shen et al.SC 2025 · 2 citations
- ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault ToleranceTong Xie, Jiawang Zhao, Zishen Wan, Zuodong Zhang et al.DAC 2025 · 4 citations
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev et al.EuroSys 2024 · 23 citations
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural NetworksSigma Jahan, Saurabhsingh Rajput, Tushar Sharma, Masud RahmanICSE 2026
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng et al.ASE 2024 · 2 citations
