MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
Zihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri, Deva Ramanan
摘要
We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild. In contrast, we aim to reconstruct dynamic human behaviors, such as repairing a bike or dancing, from a small set of sparse-view cameras with complete scene coverage (e.g. four equidistant inward-facing static cameras). We find that dense multi-view reconstruction methods struggle to adapt to this sparse-view setup due to limited overlap between viewpoints. To address these limitations, we carefully align independent monocular reconstructions of each camera to produce time- and view-consistent dynamic scene reconstructions. Extensive experiments on PanopticStudio and Ego-Exo4D demonstrate that our method achieves higher quality reconstructions than prior art, particularly when rendering novel views. Code, data, and data-processing scripts are available on https://github.com/Z1hanW/MonoFusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu 等CVPR 2026 · 被引用 23 次
- Instant4D: 4D Gaussian Splatting in MinutesZhanpeng Luo, Haoxi Ran, Li LuNeurIPS 2025 · 被引用 11 次
- 4C4D: 4 Camera 4D Gaussian SplattingJunsheng Zhou, Zhifan Yang, Liang Han, Wenyuan Zhang 等CVPR 2026 · 被引用 4 次
- FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow DerivativesQizhi Chen, Delin Qu, Junli Liu, Yiwen Tang 等AAAI 2026 · 被引用 1 次
- Contact-guided Real2Sim from Monocular Video with Planar Scene PrimitivesZihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao 等ICLR 2026
它引用的顶会 Paper40
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
相关 Paper
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim 等SIGGRAPH 2026
- SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse CamerasWeihong Pan, Xiaoyu Zhang, Zhuang Zhang, Zhichao Ye 等CVPR 2026
- FlexNeRF: Photorealistic Free-viewpoint Rendering of Moving Humans from Sparse ViewsVinoj Jayasundara, Amit Agrawal, Nicolas Heron, Abhinav Shrivastava 等CVPR 2023
- Multi-View 3D Point TrackingFrano Rajic, Haofei Xu, Marko Mihajlovic, Siyuan Li 等ICCV 2025 · 被引用 2 次
- TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view VideosJinpeng Liu, Yukang Xu, Yutong Li, Xingyu LiuCVPR 2026 · 被引用 1 次
