Lune

ICML2025Top-tier venue

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen

2025Year
49Top-tier citations

Abstract

Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited. To further enhance the efficiency of attention computation compared to SageAttention while maintaining precision, we propose SageAttention2, which utilizes significantly faster 4-bit matrix multiplication (Matmul) alongside additional precision-enhancing techniques. First, we propose to quantize matrices (Q, K) to INT4 in a hardware-friendly threadlevel granularity and quantize matrices ( P , V ) to FP8. Second, we propose a method to smooth Q, enhancing the accuracy of INT4 QK ⊤ . Third, we propose a two-level accumulation strategy for P V to enhance the accuracy of FP8 P V . The operations per second (OPS) of SageAttention2 surpass FlashAttention2 and xformers by about 3x and 4.5x. Moreover, SageAttention2 matches the speed of FlashAttention3(fp8) on the Hopper GPUs, while delivering much higher accuracy. Comprehensive experiments confirm that our approach incurs negligible end-to-end metrics loss across diverse models, including those for language, image, and video generation. The code is available at https://github.com/ thu-ml/SageAttention .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f6452a55-0ee0-4238-8146-efeb0722673d

Cited by top-tier papers49

Ask how each one uses it

Builds on36

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines