VMonarch: Efficient Video Diffusion Transformers with Structured Attention
Cheng Liang, Haoxian Chen, Liang Hou, Qi Fan, Gangshan Wu, Xin Tao, Limin Wang
Abstract
The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal attention patterns exhibited in Video DiTs can be naturally represented by the Monarch matrix. It is a class of structured matrices with flexible sparsity, enabling sub-quadratic attention via an alternating minimization algorithm. Accordingly, we propose VMonarch, a novel attention mechanism for Video DiTs that enables efficient computation over the dynamic sparse patterns with structured Monarch matrices. First, we adapt spatio-temporal Monarch factorization to explicitly capture the intra-frame and inter-frame correlations of the video data. Second, we introduce a recomputation strategy to mitigate artifacts arising from instabilities during alternating minimization of Monarch matrices. Third, we propose a novel online entropy algorithm fused into FlashAttention, enabling fast Monarch matrix updates for long sequences. Extensive experiments demonstrate that VMonarch achieves comparable or superior generation quality to full attention on VBench after minimal fine-tuning. It overcomes the attention bottleneck in Video DiTs, reduces attention FLOPs by a factor of , and achieves a speedup of over in attention computation for long videos, surpassing state-of-the-art sparse attention methods at 90% sparsity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b4db237-4797-460f-890c-b56db117b22cBuilds on50
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured AttentionCan Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee et al.NeurIPS 2025 · 5 citations
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen et al.AAAI 2026
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao et al.NeurIPS 2025 · 25 citations
- Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion TransformersYuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang et al.ICML 2026 · 5 citations
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsJianzong Wu, Liang Hou, Haotian Yang, Ye Tian et al.ICLR 2026 · 36 citations
