Move-in-2D: 2D-Conditioned Human Motion Generation
Hsin-Ping Huang, Yang Zhou, Jui-Hsien Wang, Difan Liu, Feng Liu, Ming-Hsuan Yang, Zhan Xu
Abstract
Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which restricts applications to specific motion types and global scene matching. We propose Move-in-2D, a novel approach to generate human motion sequences conditioned on a scene image, allowing for diverse motion that adapts to different scenes. Our approach utilizes a diffusion model that accepts both a scene image and text prompt as inputs, producing a motion sequence tailored to the scene. To train this model, we collect a large-scale video dataset featuring single-human activities, annotating each video with the corresponding human motion as the target output. Experiments demonstrate that our method effectively predicts human motion that aligns with the scene image after projection. Furthermore, we show that the generated motion sequence improves human motion quality in video synthesis tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3893eb0-22b0-4e61-b031-570e3dfacc38Cited by top-tier papers3
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance DecouplingHaoyu Wang, Hao Tang, Donglin Di, Zhilu Zhang et al.ICLR 2026 · 4 citations
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
- MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion GenerationWenfeng Song, Xuehan Wang, Shuai Li, Yi Chen et al.CVPR 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- Make-An-Animation: Large-Scale Text-conditional 3D Human Motion GenerationSamaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh et al.ICCV 2023 · 70 citations
- ActAnywhere: Subject-Aware Video Background GenerationBoxiao Pan, Zhan Xu, Chun-Hao Paul Huang, Krishna Kumar Singh et al.NeurIPS 2024 · 10 citations
- Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion GenerationsRuoxi Guo, Huaijin Pi, Zehong Shen, Qing Shuai et al.ICCV 2025 · 2 citations
- Putting People in Their Place: Affordance-Aware Human Insertion into ScenesSumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu et al.CVPR 2023
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
