VMoBA: Mixture-of-Block Attention for Video Diffusion Models
Jianzong Wu, Liang Hou, Haotian Yang, Ye Tian, Pengfei Wan, Di ZHANG, Yunhai Tong
摘要
The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. While various sparse attention methods have been proposed, many are designed as training-free inference accelerators or do not optimally capture the unique spatio-temporal characteristics inherent in video data when trained natively. This paper introduces Video Mixture of Block Attention (VMoBA), a novel sparse attention mechanism specifically adapted for VDMs. Motivated by an in-depth analysis of attention patterns within pre-trained video transformers, which revealed strong spatio-temporal locality, varying query importance, and head-specific concentration levels, VMoBA enhances the original MoBA framework with three key modifications: (1) a layer-wise recurrent block partition scheme (1D-2D-3D) to dynamically adapt to diverse spatio-temporal attention patterns and improve efficiency; (2) global block selection to prioritize the most salient query-key block interactions across an entire attention head; and (3) threshold-based block selection to dynamically determine the number of attended blocks based on their cumulative similarity. Extensive experiments demonstrate that VMoBA significantly accelerates the training of VDMs on longer sequences, achieving 2.92 FLOPs and 1.48 latency speedup, while attaining comparable or even superior generation quality to full attention. Furthermore, VMoBA exhibits competitive performance in training-free inference, offering 2.40 FLOPs and 1.35 latency speedup for high-res video generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li 等NeurIPS 2025 · 被引用 114 次
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo 等ICLR 2026 · 被引用 92 次
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang 等ICLR 2026 · 被引用 57 次
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video GenerationWeinan Jia, Yuning Lu, Mengqi Huang, Hualiang Wang 等ICLR 2026 · 被引用 14 次
- Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse AttentionChengtao Lv, Yumeng Shi, Yushi Huang, Ruihao Gong 等ICML 2026 · 被引用 10 次
它引用的顶会 Paper29
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
相关 Paper
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao 等NeurIPS 2025 · 被引用 25 次
- Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion TransformersYuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang 等ICML 2026 · 被引用 5 次
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan 等CVPR 2026 · 被引用 2 次
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin 等ICLR 2026
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu 等ICML 2025
