Make It Move: Controllable Image-to-Video Generation with Text Descriptions
Yaosi Hu, Chong Luo, Zhenzhong Chen
Abstract
Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Imageto-Video generation (TI2V), is proposed. With both controllable appearance and motion, TI2V aims at generating videos from a static image and a text description. The key challenges of TI2V task lie both in aligning appearance and motion from different modalities, and in handling uncertainty in text descriptions. To address these challenges, we propose a Motion Anchor-based video GEnerator (MAGE) with an innovative motion anchor (MA) structure to store appearance-motion aligned representation. To model the uncertainty and increase the diversity, it further allows the injection of explicit condition and implicit randomness. Through three-dimensional axial transformers, MA is interacted with given image to generate next frames recursively with satisfying controllability and diversity. Accompanying the new task, we build two new video-text paired datasets based on MNIST and CATER for evaluation. Experiments conducted on these datasets verify the effectiveness of MAGE and show appealing potentials of TI2V task. Datasets are available at https:// github.com/ Youncy-Hu/ MAGE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- Seer: Language Instructed Video Prediction with Latent Diffusion ModelsXianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song et al.ICLR 2024 · 57 citations
- Towards Consistent Video Editing with Text-to-Image Diffusion ModelsZicheng Zhang, Bonan Li, Xuecheng Nie, Congying Han et al.NeurIPS 2023 · 48 citations
- Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in VideosZhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing et al.ACM MM 2024 · 14 citations
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy et al.CVPR 2026 · 7 citations
- Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video GenerationAram Davtyan, Paolo FavaroAAAI 2024 · 7 citations
Builds on10
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 198 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Vid2Game: Controllable Characters Extracted from Real-World VideosOran Gafni, Lior Wolf, Yaniv TaigmanICLR 2020 · 42 citations
Related papers
- TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationXingrui Wang, Xin Li, Yaosi Hu, Hanxin Zhu et al.AAAI 2025 · 3 citations
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace et al.ICLR 2025
- MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete DiffusionOnkar Kishor Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal et al.ICLR 2025
- UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingJianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo et al.ACM MM 2025 · 6 citations
