Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
Shuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li, Jintao Zhang, Han Cai, Yujun Lin, Xiuyu Li, Chenfeng Xu, Kelly Peng, Jianfei Chen, Song Han
Abstract
Diffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention. By computing only critical tokens, sparse attention reduces computational costs and offers a promising acceleration approach. However, we identify that existing methods fail to approach optimal generation quality under the same computation budget for two reasons: (1) Inaccurate critical token identification: current methods cluster tokens based on position rather than semantics, leading to imprecise aggregated representations. (2) Excessive computation waste: critical tokens are scattered among non-critical ones, leading to wasted computation on GPUs, which are optimized for processing contiguous tokens. In this paper, we propose SVG2, a training-free framework that maximizes identification accuracy and minimizes computation waste, achieving a Pareto frontier trade-off between generation quality and efficiency. The core of SVG2 is semantic-aware permutation, which clusters and reorders tokens based on semantic similarity using k-means. This approach ensures both a precise cluster representation, improving identification accuracy, and a densified layout of critical tokens, enabling efficient computation without padding. Additionally, SVG2 integrates top-p dynamic budget control and customized kernel implementations, achieving up to 2.30x and 1.89x speedup while maintaining a PSNR of up to 30 and 26 on HunyuanVideo and Wan 2.1, respectively. Our code is open-sourced at https://github.com/svg-project/Sparse-VideoGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3d5a1be-d7da-4fe1-894a-1297e1e8806bCited by top-tier papers42
- SANA-Video: Efficient Video Generation with Block Linear Diffusion TransformerJunsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu et al.ICLR 2026 · 96 citations
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit TrainingJintao Zhang, Jia Wei, Haoxu Wang, Pengle Zhang et al.NeurIPS 2025 · 81 citations
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang et al.ICLR 2026 · 57 citations
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li et al.CVPR 2026 · 42 citations
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
Builds on44
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
Related papers
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu et al.ICML 2025
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video GenerationHaoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang et al.CVPR 2026 · 5 citations
- Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-ClusteringJiayi Luo, Jiayu Chen, Jiankun Wang, Cong Wang et al.ICML 2026 · 5 citations
- DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video GenerationJie Hu, Zixiang Gao, Yutong He, Kun YuanICML 2026
- Dynamic Sparsity in Large-Scale Video DiT TrainingXin Tan, Yuetao Chen, Yimin Jiang, Xing Chen et al.ASPLOS 2026 · 1 citation
