Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
Jie Tian, Xiaoye Qu, Zhenyi Lu, Wei Wei, Sichen Liu, Yu Cheng
Abstract
Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance of the images. However, current I2V diffusion models (I2V-DMs) often produce videos with limited motion degrees or exhibit uncontrollable motion that conflicts with the textual condition. To address these limitations, we propose a novel Extrapolating and Decoupling framework, which introduces model merging techniques to the I2V domain for the first time. Specifically, our framework consists of three separate stages: (1) Starting with a base I2V-DM, we explicitly inject the textual condition into the temporal module using a lightweight, learnable adapter and fine-tune the integrated model to improve motion controllability. (2) We introduce a training-free extrapolation strategy to amplify the dynamic range of the motion, effectively reversing the fine-tuning process to enhance the motion degree significantly. (3) With the above two-stage models excelling in motion controllability and degree, we decouple the relevant parameters associated with each type of motion ability and inject them into the base I2V-DM. Since the I2V-DM handles different levels of motion controllability and dynamics at various denoising time steps, we adjust the motion-aware parameters accordingly over time. Extensive qualitative and quantitative experiments have been conducted to demonstrate the superiority of our framework over existing methods. Code is available at https://github.com/Chuge0335/EDG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d54686c6-9152-4ff0-a9c3-0b73d7dfd41eCited by top-tier papers5
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation TasksRuibin Li, Tao Yang, Yangming Shi, Weiguo Feng et al.ICLR 2026 · 4 citations
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass GuidanceJune Suk Choi, Kyungmin Lee, Sihyun Yu, Yisol Choi et al.CVPR 2026 · 4 citations
- UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition FrameworkRui Jiang, Chongmian Wang, Xinghe Fu, Yehao Lu et al.AAAI 2026
- TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory AlignmentJin Wang, Jianxiang Lu, Guangzheng Xu, Comi Chen et al.ICML 2026
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian et al.SIGGRAPH 2024 · 66 citations
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel et al.ICLR 2026 · 13 citations
- Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion ModelMin Zhao, Hongzhou Zhu, Chendong Xiang, Kaiwen Zheng et al.NeurIPS 2024 · 33 citations
- TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationXingrui Wang, Xin Li, Yaosi Hu, Hanxin Zhu et al.AAAI 2025 · 3 citations
- VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion ModelsYabo Zhang, Yuxiang Wei, Xianhui Lin, Zheng Hui et al.AAAI 2025 · 3 citations
