Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Haiquan Qiu, Quanming Yao
Abstract
The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this progress is often hindered by notorious training instabilities. This paper provides the first mechanistic explanation for a long-standing and unresolved failure case where training with flash attention in low-precision settings leads to catastrophic loss explosion. Our indepth analysis reveals that the failure is not a random artifact but caused by two intertwined phenomena: the emergence of similar low-rank representations within the attention mechanism and the compounding effect of biased rounding errors inherent in low-precision arithmetic. We demonstrate how these factors create a vicious cycle of error accumulation that corrupts weight updates, ultimately derailing the training dynamics. To validate our findings, we introduce a minimal modification to the flash attention that mitigates the bias in rounding errors. This simple change stabilizes the training process, confirming our analysis and offering a practical solution to this persistent problem. Code is available at https: //github.com/ucker/why-low-precision-training-fails .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dcb391d-578e-41b2-b90b-a9ae7acc8e17Cited by top-tier papers2
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For DummiesGuang Liang, Jie Shao, Ningyuan Tang, Xinyao Liu et al.CVPR 2026 · 5 citations
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient ClippingGuoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng et al.ICML 2026 · 3 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Attn-QAT: 4-Bit Attention With Quantization-Aware TrainingPeiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang et al.ICML 2026 · 2 citations
- BinaryAttention: One-Bit QK-Attention for Vision and Diffusion TransformersChaodong XIAO, Zhengqiang ZHANG, Lei ZhangCVPR 2026 · 1 citation
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- Transformer Quality in Linear TimeWeizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. LeICML 2022 · 335 citations
- FlashBias: Fast Computation of Attention with BiasHaixu Wu, Minghao Guo, Yuezhou Ma, Yuanxu Sun et al.NeurIPS 2025 · 13 citations
