SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention
Jintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang, Kaiwen Zheng, Haocheng Xi, Ziteng Wang, Hongzhou Zhu, Min Zhao, Ion Stoica, Joseph E. Gonzalez, Jianfei Chen, Jun Zhu
Abstract
In Diffusion Transformer (DiT) models, particularly for video generation, attention latency is a major bottleneck due to the long sequence length and the quadratic complexity. Interestingly, we find that attention weights can be decoupled into two matrices: a small fraction of large weights with high rank and the remaining weights with very low rank. This naturally suggests applying sparse acceleration to the first part and low-rank acceleration to the second. Based on this finding, we propose SLA (Sparse-Linear Attention), a trainable attention method that fuses sparse and linear attention to accelerate diffusion models. SLA classifies attention weights into critical, marginal, and negligible, applying O(N 2 ) attention to critical weights, O(N ) attention to marginal weights, and skipping negligible ones. SLA combines these computations into a single GPU kernel and supports both forward and backward passes. With only a few fine-tuning steps using SLA, DiT models achieve a 20× reduction in attention computation, resulting in significant acceleration without loss of generation quality. Experiments show that SLA reduces attention computation by 95% without degrading end-to-end generation quality, outperforming baseline methods. In addition, we implement an efficient GPU kernel for SLA, which yields a 13.7× speedup in attention computation and a 2.2× end-to-end speedup in video generation on Wan2.1-1.3B. The code will be available at https://github.com/thu-ml/SLA . INTRODUCTION Among the operations in Transformers, attention (Vaswani et al., 2017) is the only one with quadratic computation complexity, while others mostly scale linearly with the sequence length N . In Diffusion Transformer (DiT) models (Peebles & Xie, 2022) , especially for video generation, attention becomes the primary computational bottleneck, as the sequence length typically ranges from 10K to 100K. Reducing the cost of attention is therefore critical for improving the efficiency of DiT models. Existing efficient attention methods (Zhang et al.) for DiTs fall into two main categories: (1) numerous sparse attention methods (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30b84307-6831-4b33-80fd-d73dd6255073Cited by top-tier papers8
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit TrainingJintao Zhang, Jia Wei, Haoxu Wang, Pengle Zhang et al.NeurIPS 2025 · 81 citations
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse AttentionChengtao Lv, Yumeng Shi, Yushi Huang, Ruihao Gong et al.ICML 2026 · 10 citations
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video GenerationYushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv et al.CVPR 2026 · 9 citations
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan et al.CVPR 2026 · 2 citations
Builds on26
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
Related papers
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin et al.ICLR 2026
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang et al.ICCV 2025 · 6 citations
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin et al.NeurIPS 2025 · 6 citations
- Dynamic Sparsity in Large-Scale Video DiT TrainingXin Tan, Yuetao Chen, Yimin Jiang, Xing Chen et al.ASPLOS 2026 · 1 citation
- Fast Video Generation with Sliding Tile AttentionPeiyuan Zhang, Yongqi Chen, Runlong Su, Hangliang Ding et al.ICML 2025
