MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
Zihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri, Deva Ramanan
Abstract
We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild. In contrast, we aim to reconstruct dynamic human behaviors, such as repairing a bike or dancing, from a small set of sparse-view cameras with complete scene coverage (e.g. four equidistant inward-facing static cameras). We find that dense multi-view reconstruction methods struggle to adapt to this sparse-view setup due to limited overlap between viewpoints. To address these limitations, we carefully align independent monocular reconstructions of each camera to produce time- and view-consistent dynamic scene reconstructions. Extensive experiments on PanopticStudio and Ego-Exo4D demonstrate that our method achieves higher quality reconstructions than prior art, particularly when rendering novel views. Code, data, and data-processing scripts are available on https://github.com/Z1hanW/MonoFusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e3b3132-e535-4b2d-8e8b-dd9c52d10d25Cited by top-tier papers8
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu et al.CVPR 2026 · 23 citations
- Instant4D: 4D Gaussian Splatting in MinutesZhanpeng Luo, Haoxi Ran, Li LuNeurIPS 2025 · 11 citations
- 4C4D: 4 Camera 4D Gaussian SplattingJunsheng Zhou, Zhifan Yang, Liang Han, Wenyuan Zhang et al.CVPR 2026 · 4 citations
- FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow DerivativesQizhi Chen, Delin Qu, Junli Liu, Yiwen Tang et al.AAAI 2026 · 1 citation
- Contact-guided Real2Sim from Monocular Video with Planar Scene PrimitivesZihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao et al.ICLR 2026
Builds on40
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim et al.SIGGRAPH 2026
- SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse CamerasWeihong Pan, Xiaoyu Zhang, Zhuang Zhang, Zhichao Ye et al.CVPR 2026
- FlexNeRF: Photorealistic Free-viewpoint Rendering of Moving Humans from Sparse ViewsVinoj Jayasundara, Amit Agrawal, Nicolas Heron, Abhinav Shrivastava et al.CVPR 2023
- Multi-View 3D Point TrackingFrano Rajic, Haofei Xu, Marko Mihajlovic, Siyuan Li et al.ICCV 2025 · 2 citations
- TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view VideosJinpeng Liu, Yukang Xu, Yutong Li, Xingyu LiuCVPR 2026 · 1 citation
