MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
Hsin-Ying Lee, Hanwen Jiang, Yiqun Mei, Jing Shi, Ming-Hsuan Yang, Zhixin Shu
Abstract
Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by missing secondary causal consequences. To address this, we introduce MotiMotion, a novel framework that reformulates motion control as a reasoning-then-generation problem. To encourage causally grounded and commonsense-consistent interactions, we leverage a training-free vision-language reasoner to refine image-space coordinates of primary trajectories and to hallucinate plausible secondary motions. To further improve motion naturalness, we propose a confidence-aware control scheme that modulates guidance strength, enabling the model to closely follow high-confidence plans while correcting artifacts under low-confidence inputs with its internal generative priors. To support systematic evaluation, we curate a new image-to-video benchmark, MotiBench, consisting of interaction-centric scenes where new events are triggered by motion. Both VLM-based evaluation and a human study on MotiBench demonstrate that MotiMotion produces videos with more plausible object behaviors and interaction, and is preferred over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40e2ba34-fbf2-4d2d-89ff-6c2a996f3814Builds on47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen et al.NeurIPS 2023 · 579 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
Related papers
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang et al.ICCV 2025 · 10 citations
- Motion Prompting: Controlling Video Generation with Motion TrajectoriesDaniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole et al.CVPR 2025
- VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorXindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin et al.ICCV 2025 · 8 citations
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen et al.CVPR 2026 · 8 citations
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 6 citations
