Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts
Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Chenyang Qi, Jixuan Ying, Chengfei Cai, Zhifeng Li, Heung-Yeung Shum, Wei Liu, Qifeng Chen
Abstract
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. Additionally, current I2V methods require users not only to describe the target motion but also to provide redundant detailed descriptions of frame contents. These two issues hinder the practical utilization of current I2V tools. In this paper, we propose a practical framework, named Follow-Your-Click, to achieve image animation with a simple user click (for specifying what to move) and a motion prompt (for specifying how to move). Technically, we propose the first-frame masking strategy, which significantly improves the video generation quality, and a motion-augmented module equipped with a motion prompt dataset to improve the motion prompt following abilities of our model. To further control the motion speed, we propose flow-based motion magnitude control to control the speed of target movement more precisely. Our framework has simpler yet precise user control and better generation performance than previous methods. Extensive experiments compared with 7 baselines, including both commercial tools and research methods on 8 metrics, suggest the superiority of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2a42216-757d-442d-ae92-33a4cd7c754fCited by top-tier papers27
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang et al.ICLR 2026 · 44 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlZeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang et al.ICLR 2026 · 29 citations
- DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous VariablesXiangfei Qiu, Yuhan Zhu, Zhengyu Li, Xingjian Wu et al.ICML 2026 · 22 citations
- GCGNet: Graph-Consistent Generative Network for Time Series Forecasting with Exogenous VariablesZhengyu Li, Xiangfei Qiu, Yuhan Zhu, Xingjian Wu et al.ICLR 2026 · 18 citations
Builds on31
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- Click to Move: Controlling Video Generation with Sparse MotionPierfrancesco Ardino, Marco De Nadai, Bruno Lepri, Elisa Ricci et al.ICCV 2021 · 15 citations
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin et al.CVPR 2025
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang et al.ICLR 2026
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 6 citations
- MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video GenerationJinbo Xing, Long Mai, Cusuh Ham, Jiahui Huang et al.SIGGRAPH 2025 · 11 citations
