MV-Diffusion: Motion-aware Video Diffusion Model
Zijun Deng, Xiangteng He, Yuxin Peng, Xiongwei Zhu, Lele Cheng
Abstract
In this paper, we present a Motion-aware Video Diffusion Model (MV-Diffusion) for enhancing the temporal consistency of generated videos using autoregressive diffusion models. Despite the success of diffusion models in various vision generation tasks, generating high-quality and realistic videos with coherent temporal structure remains a challenging problem. Current methods have primarily focused on capturing implicit motion features within a restricted window of RGB frames, rather than explicitly modeling the motion. To address this, we focus on improving the temporal modeling ability of the current autoregressive video diffusion approach by leveraging rich temporal trajectory information in a global context and explicitly modeling local motion trends. The main contributions of this research include: (1) a Trajectory Modeling (TM) block that enhances the model's conditioning by incorporating global motion trajectory information, (2) a Motion Trend Attention (MTA) block that utilizes a cross-attention mechanism to explicitly infer motion trends from the optical flow rather than implicitly learning from RGB input. Experimental results on three video generation tasks using four datasets show the effectiveness of our proposed MV-Diffusion, outperforming existing state-of-the-art approaches. The code is available at https://github.com/PKU-ICST-MIPL/MV-Diffusion_ACMMM2023.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52b7e9f9-82c6-4d21-bcb6-558017d09df2Cited by top-tier papers6
- Pro3D-Editor: A Progressive-Views Perspective for Consistent and Precise 3D EditingYang Zheng, Mengqi Huang, Nan Chen, Zhendong MaoNeurIPS 2025 · 11 citations
- PanoDiT: Panoramic Videos Generation with Diffusion TransformerMuyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang et al.AAAI 2025 · 6 citations
- TiP4GEN: Text to Immersive Panorama 4D Scene GenerationKe Xing, Hanwen Liang, Dejia Xu, Yuyang Yin et al.ACM MM 2025 · 2 citations
- NS-Diff: Fluid Navier-Stokes Guided Video Diffusion via Reinforcement LearningZijun Deng, Yuxin PengCVPR 2026
- Efficient Diffusion as Low Light EnhancerGuanzhou Lan, Qianli Ma, Yuqi Yang, Zhigang Wang et al.CVPR 2025
Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Trajectory attention for fine-grained video motion controlZeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang et al.ICLR 2025
- AR-Diffusion: Asynchronous Video Generation with Auto-Regressive DiffusionMingzhen Sun, Weining Wang, Gen Li, Jiawei Liu et al.CVPR 2025
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- MoStGAN-V: Video Generation with Temporal Motion StylesXiaoqian Shen, Xiang Li, Mohamed ElhoseinyCVPR 2023
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian et al.SIGGRAPH 2024 · 66 citations
