Reconstructing People, Places, and Cameras
Lea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi, Jitendra Malik, Angjoo Kanazawa
Abstract
Humans and Structure from Motion (HSfM). We propose a method for the joint reconstruction of humans, scene point clouds, and cameras from an uncalibrated, sparse set of images depicting people. By explicitly incorporating humans into the traditional Structure from Motion (SfM) framework through 2D human keypoint correspondences and leveraging robust initialization from an off-theshelf model for scene and camera reconstruction, our approach demonstrates that integrating these three elements-people, scenes, and cameras-synergistically improves the reconstruction accuracy of each component. Unlike prior work in SfM and human pose estimation, our method reconstructs metric-scale scene point clouds and camera parameters, informed by human mesh predictions, while situating human meshes in coherent world coordinates consistent with the surrounding environment without any explicit contact constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76fc78a1-cd88-4d58-a823-db336b31694dCited by top-tier papers12
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- Joint Optimization for 4D Human-Scene Reconstruction in the WildZhizheng Liu, Joe Lin, Wayne Wu, Bolei ZhouICLR 2026 · 33 citations
- UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward PassMengfei Li, Peng Li, Zheng Zhang, Jiahao Lu et al.CVPR 2026 · 7 citations
- OnlineHMR: Video-based Online World-Grounded Human Mesh RecoveryYiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang et al.CVPR 2026 · 5 citations
- HAMSt3R: Human-Aware Multi-View Stereo 3D ReconstructionSara Rojas, Matthieu Armando, Bernard Ghanem, Philippe Weinzaepfel et al.ICCV 2025 · 3 citations
Builds on26
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-CompareYuanlu Xu, Song-Chun Zhu, Tony TungICCV 2019 · 204 citations
Related papers
- Synergistic Global-Space Camera and Human Reconstruction from VideosYizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu et al.CVPR 2024 · 2 citations
- Detector-Free Structure from MotionXingyi He, Jiaming Sun, Yifan Wang, Sida Peng et al.CVPR 2024
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang et al.NeurIPS 2022 · 24 citations
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
- Wide-Baseline Multi-Camera Calibration Using Person Re-IdentificationYan Xu, Yu-Jhe Li, Xinshuo Weng, Kris KitaniCVPR 2021
