Learning Human Dynamics in Autonomous Driving Scenarios
Jingbo Wang, Ye Yuan, Zhengyi Luo, Kevin Xie, Dahua Lin, Umar Iqbal, Sanja Fidler, Sameh Khamis
Abstract
Simulation has emerged as an indispensable tool for scaling and accelerating the development of self-driving systems. A critical aspect of this is simulating realistic and diverse human behavior and intent. In this work, we propose a holistic framework for learning physically plausible human dynamics from real driving scenarios, narrowing the gap between real and simulated human behavior in safety-critical applications. We show that state-of-the-art methods underperform in driving scenarios where video data is recorded from moving vehicles, and humans are frequently partially or fully occluded. Furthermore, existing methods often disregard the global scene where humans are situated, resulting in various motion artifacts like foot sliding, floating, or ground penetration. To address this challenge, we propose an approach that incorporates physics with a reinforcement learning-based motion controller to learn human dynamics for driving scenarios. Our framework can simulate physically plausible human dynamics that accurately match observed human motions and infill motions for occluded body parts, while improving the physical plausibility of the entire motion sequence. Experiments on the challenging Waymo Open Dataset show that our method outperforms state-of-the-art motion capture approaches significantly in recovering high-quality, physically plausible, and scene-aware human dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- A Plug-And-Play Physical Motion Restoration Approach for In-The-Wild High-Difficulty MotionsYouliang Zhang, Ronghui Li, Yachao Zhang, Liang Pan et al.ICCV 2025 · 4 citations
- Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction RecognitionShiyu Xuan, Dongkai Wang, Zechao Li, Jinhui TangICLR 2026 · 2 citations
- Towards Unstructured Unlabeled Optical Mocap: A Video Helps!Nicholas Milef, John Keyser, Shu KongSIGGRAPH 2024 · 2 citations
- FusionSAM: Visual Multi-Modal Learning with Segment Anything ModelDaixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang et al.KDD 2025 · 2 citations
- ImDy: Human Inverse Dynamics from Imitated ObservationsXinpeng Liu, Junxuan Liang, Zili Lin, Haowen Hou et al.ICLR 2025 · 1 citation
Builds on26
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
Related papers
- QuestEnvSim: Environment-Aware Simulated Motion Tracking from Sparse SensorsSunmin Lee, Sebastian Starke, Yuting Ye, Jungdam Won et al.SIGGRAPH 2023 · 31 citations
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- Learning Motion Priors for 4D Human Body Capture in 3D ScenesSiwei Zhang, Yan Zhang, Federica Bogo, Marc Pollefeys et al.ICCV 2021 · 117 citations
- Physics-based Scene Layout Generation from Human MotionJianan Li, Tao Huang, Qingxu Zhu, Tien-Tsin WongSIGGRAPH 2024 · 5 citations
- MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondShenghao Ren, Yi Lu, Jiayi Huang, Jiayi Zhao et al.CVPR 2025
