Decoupling Human and Camera Motion from Videos in the Wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, Angjoo Kanazawa
Abstract
Human Motion in the World Frame Figure 1. 4D Reconstruction of People from Videos in-the-Wild. We present SLAHMR: Simultaneous Localization And Human Mesh Recovery, a method that given a video of moving people (top), recovers the global trajectories of all people and cameras in the world coordinate space (bottom). We combine geometric insights, which determine relative camera motion, with learned human motion priors, which constrain a person's plausible displacement between frames, to position the people and cameras in the shared world frame through time. Our method can recover the global trajectories of all detected people from in-the-wild videos with uncontrolled camera and human motion. Please see the project page to see the full video results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers69
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the WildManuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen et al.ICCV 2023 · 94 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
- WHAM: Reconstructing World-Grounded Humans with Accurate 3D MotionSoyong Shin, Juyong Kim, Eni Halilaj, Michael J. BlackCVPR 2024 · 66 citations
- Deformable Neural Radiance Fields using RGB and Event CamerasQi Ma, Danda Pani Paudel, Ajad Chhatkuli, Luc Van GoolICCV 2023 · 43 citations
Builds on30
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
Related papers
- Synergistic Global-Space Camera and Human Reconstruction from VideosYizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu et al.CVPR 2024 · 2 citations
- Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic CameraZhengdi Yu, Stefanos Zafeiriou, Tolga BirdalCVPR 2025
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera DecouplingFengyuan Yang, Tanuj Sur, Tze Ho Elden Tse, Angela YaoCVPR 2026
