VMoBA: Mixture-of-Block Attention for Video Diffusion Models
Jianzong Wu, Liang Hou, Haotian Yang, Ye Tian, Pengfei Wan, Di ZHANG, Yunhai Tong
Abstract
The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. While various sparse attention methods have been proposed, many are designed as training-free inference accelerators or do not optimally capture the unique spatio-temporal characteristics inherent in video data when trained natively. This paper introduces Video Mixture of Block Attention (VMoBA), a novel sparse attention mechanism specifically adapted for VDMs. Motivated by an in-depth analysis of attention patterns within pre-trained video transformers, which revealed strong spatio-temporal locality, varying query importance, and head-specific concentration levels, VMoBA enhances the original MoBA framework with three key modifications: (1) a layer-wise recurrent block partition scheme (1D-2D-3D) to dynamically adapt to diverse spatio-temporal attention patterns and improve efficiency; (2) global block selection to prioritize the most salient query-key block interactions across an entire attention head; and (3) threshold-based block selection to dynamically determine the number of attended blocks based on their cumulative similarity. Extensive experiments demonstrate that VMoBA significantly accelerates the training of VDMs on longer sequences, achieving 2.92 FLOPs and 1.48 latency speedup, while attaining comparable or even superior generation quality to full attention. Furthermore, VMoBA exhibits competitive performance in training-free inference, offering 2.40 FLOPs and 1.35 latency speedup for high-res video generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b2f24da-f9d9-48c7-9343-29b170aa6b77Cited by top-tier papers14
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li et al.NeurIPS 2025 · 114 citations
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo et al.ICLR 2026 · 92 citations
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang et al.ICLR 2026 · 57 citations
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video GenerationWeinan Jia, Yuning Lu, Mengqi Huang, Hualiang Wang et al.ICLR 2026 · 14 citations
- Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse AttentionChengtao Lv, Yumeng Shi, Yushi Huang, Ruihao Gong et al.ICML 2026 · 10 citations
Builds on29
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
Related papers
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao et al.NeurIPS 2025 · 25 citations
- Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion TransformersYuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang et al.ICML 2026 · 5 citations
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan et al.CVPR 2026 · 2 citations
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin et al.ICLR 2026
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu et al.ICML 2025
