4D Human-Scene Reconstruction from Low-Overlap Captures
Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park
Abstract
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement. Project page: https://sisyphm.github.io/studiorecon-page/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
- How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion ModelsPascal Chang, Jingwei Tang, Markus Gross, Vinicius C. AzevedoICLR 2024 · 44 citations
- Recammaster: Camera-Controlled Generative Rendering From a Single VideoJianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang et al.ICCV 2025 · 33 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
Related papers
- Diffuman4D: 4D Consistent Human View Synthesis From Sparse-View Videos With Spatio-Temporal Diffusion ModelsYudong Jin, Sida Peng, Xuan Wang, Tao Xie et al.ICCV 2025 · 5 citations
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsMichal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang et al.NeurIPS 2025 · 5 citations
- MonoFusion: Sparse-View 4D Reconstruction via Monocular FusionZihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri et al.ICCV 2025 · 5 citations
- SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse CamerasWeihong Pan, Xiaoyu Zhang, Zhuang Zhang, Zhichao Ye et al.CVPR 2026
- 4C4D: 4 Camera 4D Gaussian SplattingJunsheng Zhou, Zhifan Yang, Liang Han, Wenyuan Zhang et al.CVPR 2026 · 4 citations
