ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
Yuhang Liang, Xinyi Li, Jie Ren, Ang Li, Bo Fang, Jieyang Chen
摘要
Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, AT-TNChecker reduces recovery overhead by up to 49×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant AttentionHuangliang Dai, Shixun Wu, Jiajun Huang, Zizhe Jian 等SC 2025 · 被引用 12 次
- Demystifying the Resilience of Large Language Model Inference: An End-to-End PerspectiveYu Sun, Zachary Coalson, Shiyang Chen, Hang Liu 等SC 2025 · 被引用 9 次
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park 等ICLR 2026
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen 等OSDI 2026
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsYi He, Mike Hutton, Steven Chan, Robert De Gruijl 等ISCA 2023 · 被引用 52 次
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng 等NSDI 2025 · 被引用 46 次
- Improving Energy Saving of One-Sided Matrix Decompositions on CPU-GPU Heterogeneous SystemsJieyang Chen, Xin Liang, Kai Zhao, Hadi Zamani Sabzi 等PPoPP 2023 · 被引用 4 次
相关 Paper
- Exploring and Mitigating Failure Behavior of Large Language Model Training Workloads in HPC SystemsPengfei Yu, Jingjing Gu, Hao Han, Dazhong Shen 等SC 2025 · 被引用 2 次
- ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault ToleranceTong Xie, Jiawang Zhao, Zishen Wan, Zuodong Zhang 等DAC 2025 · 被引用 4 次
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural NetworksSigma Jahan, Saurabhsingh Rajput, Tushar Sharma, Masud RahmanICSE 2026
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng 等ASE 2024 · 被引用 2 次
