Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Haiquan Qiu, Quanming Yao
摘要
The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this progress is often hindered by notorious training instabilities. This paper provides the first mechanistic explanation for a long-standing and unresolved failure case where training with flash attention in low-precision settings leads to catastrophic loss explosion. Our indepth analysis reveals that the failure is not a random artifact but caused by two intertwined phenomena: the emergence of similar low-rank representations within the attention mechanism and the compounding effect of biased rounding errors inherent in low-precision arithmetic. We demonstrate how these factors create a vicious cycle of error accumulation that corrupts weight updates, ultimately derailing the training dynamics. To validate our findings, we introduce a minimal modification to the flash attention that mitigates the bias in rounding errors. This simple change stabilizes the training process, confirming our analysis and offering a practical solution to this persistent problem. Code is available at https: //github.com/ucker/why-low-precision-training-fails .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For DummiesGuang Liang, Jie Shao, Ningyuan Tang, Xinyao Liu 等CVPR 2026 · 被引用 5 次
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient ClippingGuoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
相关 Paper
- Attn-QAT: 4-Bit Attention With Quantization-Aware TrainingPeiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang 等ICML 2026 · 被引用 2 次
- BinaryAttention: One-Bit QK-Attention for Vision and Diffusion TransformersChaodong XIAO, Zhengqiang ZHANG, Lei ZhangCVPR 2026 · 被引用 1 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- Transformer Quality in Linear TimeWeizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. LeICML 2022 · 被引用 335 次
- FlashBias: Fast Computation of Attention with BiasHaixu Wu, Minghao Guo, Yuezhou Ma, Yuanxu Sun 等NeurIPS 2025 · 被引用 13 次
