HAMSt3R: Human-Aware Multi-View Stereo 3D Reconstruction
Sara Rojas, Matthieu Armando, Bernard Ghanem, Philippe Weinzaepfel, Vincent Leroy, Grégory Rogez
Abstract
Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly predicting dense scene geometry, they are primarily trained on outdoor scenes with static environments and struggle to handle human-centric scenarios. In this work, we introduce HAMSt3R, an extension of MASt3R for joint human and scene 3D reconstruction from sparse, uncalibrated multi-view images. First, we exploit DUNE, a strong image encoder obtained by distilling, among others, the encoders from MASt3R and from a state-of-the-art Human Mesh Recovery (HMR) model, multi-HMR, for a better understanding of scene geometry and human bodies. Our method then incorporates additional network heads to segment people, estimate dense correspondences via DensePose, and predict depth in human-centric environments, enabling a more comprehensive 3D reconstruction. By leveraging the outputs of our different heads, HAMSt3R produces a dense point map enriched with human semantic information in 3D. Unlike existing methods that rely on complex optimization pipelines, our approach is fully feed-forward and efficient, making it suitable for real-world applications. We evaluate our model on EgoHumans and EgoExo4D, two challenging benchmarks con taining diverse human-centric scenarios. Additionally, we validate its generalization to traditional multi-view stereo and multi-view pose regression tasks. Our results demonstrate that our method can reconstruct humans effectively while preserving strong performance in general 3D reconstruction tasks, bridging the gap between human and scene understanding in 3D vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ba92d6a-7715-4cae-a6c3-541d5a530c2cCited by top-tier papers4
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward PassMengfei Li, Peng Li, Zheng Zhang, Jiahao Lu et al.CVPR 2026 · 7 citations
- EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied AgentsWenjia Wang, Liang Pan, Huaijin Pi, Yuke Lou et al.CVPR 2026 · 2 citations
- RHINO: Reconstructing Human Interactions with Novel Objects from Monocular VideosLixin Xue, Chengwei Zheng, Georgios Paschalidis, Chen Guo et al.CVPR 2026
Builds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- PanSt3R: Multi-View Consistent Panoptic SegmentationLojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld et al.ICCV 2025 · 5 citations
- MUSt3R: Multi-view Network for Stereo 3D ReconstructionYohann Cabon, Lucas Stoffl, Leonid Antsfeld, Gabriela Csurka et al.CVPR 2025
- MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsZhenggang Tang, Yuchen Fan, Dilin Wang, Hongyu Xu et al.CVPR 2025
- MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular ImagesChentao Song, He Zhang, Haolei Yuan, Haozhe Lin et al.CVPR 2026 · 5 citations
