Click to Move: Controlling Video Generation with Sparse Motion
Pierfrancesco Ardino, Marco De Nadai, Bruno Lepri, Elisa Ricci, Stéphane Lathuilière
Abstract
This paper introduces Click to Move (C2M), a novel framework for video generation where the user can control the motion of the synthesized video through mouse clicks specifying simple object trajectories of the key objects in the scene. Our model receives as input an initial frame, its corresponding segmentation map and the sparse motion vectors encoding the input provided by the user. It outputs a plausible video sequence starting from the given frame and with a motion that is consistent with user input. Notably, our proposed deep architecture incorporates a Graph Convolution Network (GCN) modelling the movements of all the objects in the scene in a holistic manner and effectively combining the sparse user motion information and image features. Experimental results show that C2M outperforms existing methods on two publicly available datasets, thus demonstrating the effectiveness of our GCN framework at modelling object interactions. The source code is publicly available at https: //github.com/PierfrancescoArdino/C2M .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Point Prompting: Counterfactual Tracking with Video Diffusion ModelsAyush Shrivastava, Sanyam Mehta, Daniel Geng, Andrew OwensICLR 2026 · 5 citations
- DragEntity: Trajectory Guided Video Generation using Entity and Positional RelationshipsZhang Wan, Sheng Tang, Jiawei Wei, Ruize Zhang et al.ACM MM 2024 · 4 citations
- Zero-Shot Controllable Image-to-Video Animation via Motion DecompositionShoubin Yu, Jacob Zhiyuan Fang, Jian Zheng, Gunnar A. Sigurdsson et al.ACM MM 2024 · 3 citations
- Motion Prompting: Controlling Video Generation with Motion TrajectoriesDaniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole et al.CVPR 2025
- MotiMotion: Motion-Controlled Video Generation with Visual ReasoningHsin-Ying Lee, Hanwen Jiang, Yiqun Mei, Jing Shi et al.ICML 2026
Builds on12
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Region Normalization for Image InpaintingTao Yu, Zongyu Guo, Xin Jin, Shilin Wu et al.AAAI 2020 · 204 citations
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
Related papers
- Follow-Your-Click: Open-domain Regional Image Animation via Motion PromptsYue Ma, Yingqing He, Hongfa Wang, Andong Wang et al.AAAI 2025 · 57 citations
- MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video GenerationJinbo Xing, Long Mai, Cusuh Ham, Jiahui Huang et al.SIGGRAPH 2025 · 11 citations
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRuihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang et al.NeurIPS 2025 · 50 citations
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang et al.ICCV 2025 · 10 citations
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin et al.CVPR 2025
