Contact-aware Human Motion Forecasting
Wei Mao, Miaomiao Liu, Richard I. Hartley, Mathieu Salzmann
Abstract
In this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challenge of this task is to ensure consistency between the human and the scene, accounting for human-scene interactions. Previous attempts to do so model such interactions only implicitly, and thus tend to produce artifacts such as "ghost motion" because of the lack of explicit constraints between the local poses and the global motion. Here, by contrast, we propose to explicitly model the human-scene contacts. To this end, we introduce distance-based contact maps that capture the contact relationships between every joint and every 3D scene point at each time instant. We then develop a two-stage pipeline that first predicts the future contact maps from the past ones and the scene point cloud, and then forecasts the future human poses by conditioning them on the predicted contact maps. During training, we explicitly encourage consistency between the global motion and the local poses via a prior defined using the contact maps and future poses. Our approach outperforms the state-of-the-art human motion forecasting and human synthesis methods on both synthetic and real datasets. Our code is available at https://github.com/wei-mao-2019/ContAwareMotionPred . Recently, a few works [8, 6] have started to incorporate scene context in motion forecasting. In particular, Corona et al. [8] introduced a semantic-graph model that extracts a joint embedding of the human pose and an object of interest, such as a cup. This method, however, is ill-suited to model interactions with the whole scene itself, for example the floor or stairs that the person touches while walking. In [6], Cao et al. proposed a multi-stage pipeline that breaks down the motion forecasting into three sub-tasks: predicting a 2D goal, planning a 2D and 3D path, forecasting the 3D poses
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f659aeeb-ec78-47b1-937f-74e032cc2b77Cited by top-tier papers19
- A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Chen ChenNeurIPS 2023 · 40 citations
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li et al.CVPR 2024 · 38 citations
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu et al.CVPR 2024 · 12 citations
- Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion PredictionZhenyu Lou, Qiongjie Cui, Tuo Wang, Zhenbo Song et al.NeurIPS 2024 · 10 citations
- Decoupled Generative Modeling for Human-Object Interaction SynthesisHwanhee Jung, Seunggwan Lee, Jeongyoon Yoon, SeungHyeon Kim et al.CVPR 2026 · 4 citations
Builds on13
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 242 citations
- Stochastic Scene-Aware Motion PredictionMohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito et al.ICCV 2021 · 240 citations
- Structured Prediction Helps 3D Human Motion ModellingEmre Aksan, Manuel Kaufmann, Otmar HilligesICCV 2019 · 204 citations
Related papers
- InterPhys: Physics-aware Human Motion Synthesis in a Dynamic SceneChaoyue Xing, Wei Mao, Miaomiao LiuCVPR 2026 · 1 citation
- Scene-Aware Generative Network for Human Motion SynthesisJingbo Wang, Sijie Yan, Bo Dai, Dahua LinCVPR 2021
- Multimodal Sense-Informed Forecasting of 3D Human MotionsZhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang et al.CVPR 2024 · 8 citations
- Task-Oriented Human Grasp Synthesis via Context- and Task-Aware DiffusersAn-Lun Liu, Yu-Wei Chao, Yi-Ting ChenICCV 2025 · 1 citation
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D ScenesJiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu et al.CVPR 2021
