SC2025Top-tier venue
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
Huangliang Dai, Shixun Wu, Jiajun Huang, Zizhe Jian, Yue Zhu, Haiyang Hu, Zizhong Chen
Abstract
Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in largescale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for Transformers are designed at the operation level without architectural optimization, leading to significant computational and memory overhead, which in turn reduces protection efficiency and limits scalability to larger models. In this paper, we implement module-level protection for Transformers by treating the operations within the attention module as a single kernel and applying end-to-end fault tolerance 1 . This method provides unified protection across multi-step computations, while achieving comprehensive coverage of potential errors in the nonlinear computations 2 . For linear modules, we design a strided algorithm-based fault tolerance (ABFT) that avoids inter-thread communication 3 . Experimental results show that our end-to-end fault tolerance achieves up to 7.56× speedup over traditional methods with an average fault tolerance overhead of 13.9%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d20c6bb-c531-414a-98f7-5e0ea7311c26Cited by top-tier papers2
- TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPUShixun Wu, Yujia Zhai, Huangliang Dai, Yue Zhu et al.SC 2025 · 5 citations
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun et al.PPoPP 2026
Builds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model TrainingYuhang Liang, Xinyi Li, Jie Ren, Ang Li et al.PPoPP 2025 · 10 citations
- TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUsShixun Wu, Yujia Zhai, Jinyang Liu, Jiajun Huang et al.PPoPP 2025 · 7 citations
- Arithmetic-intensity-guided fault tolerance for neural network inference on GPUsJack Kosaian, K. V. RashmiSC 2021 · 51 citations
- SAVE: Software-Implemented Fault Tolerance for Model Inference against GPU Memory Bit FlipsWenxin Zheng, Bin Xu, Jinyu Gu, Haibo ChenUSENIX ATC 2025 · 8 citations
- FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M SparsityManting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao et al.DAC 2024 · 10 citations
