M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao, Yinlong Qian, Ling Chen, Yunchao Wei, Lin Ma
摘要
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when employing Transformers, which incur quadratic complexity in sequence processing and thus limit practical applications. Recent advancements in linear-time sequence modeling, particularly the Mamba architecture, offer a more efficient alternative. Nevertheless, its plain design limits its direct applicability to multimodal and spatiotemporal video generation tasks. To address these challenges, we introduce M4V, a multimodal Mamba framework for efficient textto-video generation. Specifically, a MultiModal diffusion Mamba (MM-DiM) block is designed within the framework to enable seamless integration of multimodal information and spatiotemporal modeling. In detail, we introduce a novel multimodal token re-composition design, which employs a bidirectional scheme for multimodal information integration through simple token arrangement, along with visual registers to enhance spatial-temporal consistency. As a result, the MM-DiM blocks in M4V reduce FLOPs by 45% compared with the attention-based alternative when generating videos at 768×1280 resolution. Additionally, several training strategies are explored in this work to provide a better understanding of training text-to-video models using only publicly available datasets. Extensive experiments on text-to-video benchmarks demonstrate M4V's ability to produce high-quality videos while significantly lowering computational costs. Project page is available at https: //huangjch526.github.io/M4V_project/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu 等NeurIPS 2025 · 被引用 228 次
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion TransformerMohsen Ghafoorian, Denis Korzhenkov, Amirhossein HabibianCVPR 2026 · 被引用 13 次
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen 等NeurIPS 2025 · 被引用 9 次
- ReHyAt: Recurrent Hybrid Attention for Video Diffusion TransformersMohsen Ghafoorian, Amirhossein HabibianCVPR 2026 · 被引用 5 次
- VMonarch: Efficient Video Diffusion Transformers with Structured AttentionCheng Liang, Haoxian Chen, Liang Hou, Qi Fan 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational ComplexityHongjie Wang, Chih-Yao Ma, Yen-Cheng Liu, Ji Hou 等CVPR 2025
- VAMBA: Understanding Hour-Long Videos with Hybrid Mamba-TransformersWeiming Ren, Wentao Ma, Huan Yang, Cong Wei 等ICCV 2025 · 被引用 2 次
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video UnderstandingBoshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju 等CVPR 2026 · 被引用 5 次
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
- End-to-End Multi-Modal Diffusion MambaChunhao Lu, Qiang Lu, Meichen Dong, Jake LuoICCV 2025 · 被引用 1 次
