Stochastic Scene-Aware Motion Prediction
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, Michael J. Black
Abstract
A long-standing goal in computer vision is to capture, model, and realistically synthesize human behavior. Specifically, by learning from data, our goal is to enable virtual humans to navigate within cluttered indoor scenes and naturally interact with objects. Such embodied behavior has applications in virtual reality, computer games, and robotics, while synthesized behavior can be used as training data. The problem is challenging because real human motion is diverse and adapts to the scene. For example, a person can sit or lie on a sofa in many places and with varying styles. We must model this diversity to synthesize virtual humans that realistically perform human-scene interactions. We present a novel data-driven, stochastic motion synthesis method that models different styles of performing a given action with a target object. Our Scene-Aware Motion Prediction method (SAMP) generalizes to target objects of various geometries while enabling the character to navigate in cluttered scenes. To train SAMP, we collected MoCap data covering various sitting, lying down, walking, and running styles. We demonstrate SAMP on complex indoor scenes and achieve superior performance than existing solutions. Code and data are available for research at https://samp.is.tue.mpg.de .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4439088-06e4-48da-b4c4-56599a36a773Cited by top-tier papers95
- OmniControl: Control Any Joint at Any Time for Human Motion GenerationYiming Xie, Varun Jampani, Lei Zhong, Deqing Sun et al.ICLR 2024 · 228 citations
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- BEHAVE: Dataset and Method for Tracking Human Object InteractionsBharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu et al.CVPR 2022 · 144 citations
- Synthesizing Diverse Human Motions in 3D Indoor ScenesKaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler et al.ICCV 2023 · 116 citations
Builds on13
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Character controllers using motion VAEsHung Yu Ling, Fabio Zinno, George Cheng, Michiel van de PanneSIGGRAPH 2020 · 261 citations
- Diverse Trajectory Forecasting with Determinantal Point ProcessesYe Yuan, Kris M. KitaniICLR 2020 · 149 citations
- Catch & Carry: reusable neural controllers for vision-guided whole-body tasksJosh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa et al.SIGGRAPH 2020 · 103 citations
Related papers
- Synthesizing Physical Character-Scene InteractionsMohamed Hassan, Yunrong Guo, Tingwu Wang, Michael J. Black et al.SIGGRAPH 2023 · 60 citations
- Towards Diverse and Natural Scene-aware 3D Human Motion SynthesisJingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan et al.CVPR 2022 · 74 citations
- Learning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking ReferencesYitang Li, Mingxian Lin, Zhuo Lin, Yipeng Deng et al.CVPR 2025
- MIME: Human-Aware 3D Scene GenerationHongwei Yi, Chun-Hao P. Huang, Shashank Tripathi, Lea Hering et al.CVPR 2023
- Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D EnvironmentsJiye Lee, Hanbyul JooICCV 2023 · 55 citations
