VORTA: Efficient Video Diffusion via Routing Sparse Attention
Wenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao, Zhao Jin, Shunyu Liu, Dacheng Tao
Abstract
Video diffusion transformers have achieved remarkable progress in high-quality video generation, but remain computationally expensive due to the quadratic complexity of attention over high-dimensional video sequences. Recent acceleration methods enhance the efficiency by exploiting the local sparsity of attention scores; yet they often struggle with accelerating the long-range computation. To address this problem, we propose VORTA, an acceleration framework with two novel components: 1) a sparse attention mechanism that efficiently captures long-range dependencies, and 2) a routing strategy that adaptively replaces full 3D attention with specialized sparse attention variants. VORTA achieves an end-to-end speedup without loss of quality on VBench. Furthermore, it can seamlessly integrate with various other acceleration methods, such as model caching and step distillation, reaching up to speedup with negligible performance degradation. VORTA demonstrates its efficiency and enhances the practicality of video diffusion transformers in real-world settings. Codes and weights are available at https://github.com/wenhao728/VORTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li et al.NeurIPS 2025 · 114 citations
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video GenerationYushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv et al.CVPR 2026 · 9 citations
- RegionE: Adaptive Region-Aware Generation for Efficient Image EditingPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen et al.ICLR 2026 · 5 citations
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan et al.CVPR 2026 · 2 citations
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao et al.NeurIPS 2025 · 2 citations
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Veda: Scalable Video Diffusion via Distilled Sparse AttentionShihao Han, Hao Yang, Xiaofeng Mei, Xinting Hu et al.ICML 2026 · 1 citation
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin et al.ICLR 2026
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsJianzong Wu, Liang Hou, Haotian Yang, Ye Tian et al.ICLR 2026 · 36 citations
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context InferenceAnmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin et al.CVPR 2026 · 4 citations
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen et al.AAAI 2026
