RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao, Lei Sun, Yichen Qian, Weihua Chen, Fan Wang
Abstract
Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background video, human trajectory, and action patterns. In this paper, we propose a decomposed human motion control and video generation framework that explicitly decouples motion from appearance, subject from background, and action from trajectory, enabling flexible mix-and-match composition of these elements. Concretely, we first build a ground-aware 3D world coordinate system and perform motion editing directly in the 3D space. Trajectory control is implemented by unprojecting edited 2D trajectories into 3D with focal-length calibration and coordinate transformation, followed by speed alignment and orientation adjustment; actions are supplied by a motion bank or generated via text-to-motion methods. Then, based on modern text-to-video diffusion transformer models, we inject the subject as tokens for full attention, concatenate the background along the channel dimension, and add motion (trajectory and action) control signals by addition. Such a design opens up the possibility for us to generate realistic videos of anyone doing anything anywhere. Extensive experiments on benchmark datasets and real-world cases demonstrate that our method achieves state-of-the-art performance on both element-wise controllability and overall video quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e2d2fbc-99f6-47ae-862b-4eb575c5815bCited by top-tier papers4
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance DecouplingHaoyu Wang, Hao Tang, Donglin Di, Zhilu Zhang et al.ICLR 2026 · 4 citations
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D ScenesYu Li, Menghan Xia, Gongye Liu, Jianhong Bai et al.ICLR 2026 · 3 citations
- ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video GenerationOmar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque et al.SIGGRAPH 2026
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
Related papers
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace et al.ICLR 2025
- Frame In-N-Out: Unbounded Controllable Image-to-Video GenerationBoyang Wang, Xuweiyi Chen, Matheus Gadelha, Zezhou ChengNeurIPS 2025 · 9 citations
- Compositional 3D-aware Video Generation with LLM DirectorHanxin Zhu, Tianyu He, Anni Tang, Junliang Guo et al.NeurIPS 2024 · 19 citations
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang et al.AAAI 2025 · 1 citation
- ActAnywhere: Subject-Aware Video Background GenerationBoxiao Pan, Zhan Xu, Chun-Hao Paul Huang, Krishna Kumar Singh et al.NeurIPS 2024 · 10 citations
