UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
Mengfei Li, Peng Li, Zheng Zhang, Jiahao Lu, Chengfeng Zhao, Wei Xue, Qifeng Liu, Sida Peng, Wenxiao ZHANG, Wenhan Luo, Yuan Liu, Yike Guo
Abstract
We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain gap, leading to poor generalization, low-fidelity human geometry, and poor alignment on in-the-wild videos. To address this, we propose an innovative training paradigm that effectively leverages unlabeled in-the-wild data. Our framework bridges strong, disparate priors from scene reconstruction and HMR, and is trained with two core components: (1) a robust distillation strategy to refine human surface details by distilling high-frequency details from an expert depth model, and (2) a two-stage supervision scheme, which first learns coarse localization on synthetic data, then fine-tunes on real data by directly optimizing the geometric correspondence between the SMPL mesh and the human point cloud. This approach enables our feed-forward model to jointly recover high-fidelity scene geometry, human point clouds, camera parameters, and coherent, metric-scale SMPL bodies, all in a single forward pass. Extensive experiments demonstrate that our model achieves state-of-the-art performance on human-centric scene reconstruction and delivers highly competitive results on global human motion estimation, comparing favorably against both optimization-based frameworks and HMR-only methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 928b244c-f962-484d-b628-dfee7f91187fBuilds on44
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
Related papers
- Synergistic Global-Space Camera and Human Reconstruction from VideosYizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu et al.CVPR 2024 · 2 citations
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- Reconstructing People, Places, and CamerasLea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi et al.CVPR 2025
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu et al.CVPR 2020
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
