SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
Jintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu, Jianfei Chen
摘要
The transformer architecture predominates across various models. As the heart of the transformer, attention has a computational complexity of O(N 2 ), compared to O(N ) for linear transformations. When handling large sequence lengths, attention becomes the primary time-consuming component. Although quantization has proven to be an effective method for accelerating model inference, existing quantization methods primarily focus on optimizing the linear layer. In response, we first analyze the feasibility of quantization in attention detailedly. Following that, we propose SageAttention, a highly efficient and accurate quantization method for attention. The OPS (operations per second) of our approach outperforms FlashAttention2 and xformers by about 2.1x and 2.7x, respectively. SageAttention also achieves superior accuracy performance over FlashAt-tention3. Comprehensive experiments confirm that our approach incurs almost no end-to-end metrics loss across diverse models-including those for large language processing, image generation, and video generation. The code is available at https://github.com/thu-ml/SageAttention .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper70
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein 等NeurIPS 2025 · 被引用 132 次
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu 等ICML 2026 · 被引用 108 次
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo 等ICLR 2026 · 被引用 92 次
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang 等NeurIPS 2025 · 被引用 89 次
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit TrainingJintao Zhang, Jia Wei, Haoxu Wang, Pengle Zhang 等NeurIPS 2025 · 被引用 81 次
它引用的顶会 Paper41
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 QuantizationJintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei 等ICML 2025
- SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model InferenceJintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei 等ICML 2025
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran 等ACM MM 2025 · 被引用 3 次
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured AttentionCan Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee 等NeurIPS 2025 · 被引用 5 次
- Scaling Attention via Feature SparsityYan Xie, Tiansheng Wen, Tangda Huang, Bo Chen 等ICLR 2026 · 被引用 3 次
