O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model
Yubin Hu, Sheng Ye, Wang Zhao, Matthieu Lin, Yuze He, Yu-Hui Wen, Ying He, Yong-Jin Liu
摘要
Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction From Cluttered ScenesZesong Yang, Bangbang Yang, Liyuan Cui, Yuewen Ma 等ICCV 2025 · 被引用 2 次
- MIDGArD: Modular Interpretable Diffusion over Graphs for Articulated DesignsQuentin Leboutet, Nina Wiedemann, Zhipeng Cai, Michael Paulitsch 等NeurIPS 2024 · 被引用 2 次
- DecoupledGaussian: Object-Scene Decoupling for Physics-Based InteractionMiaowei Wang, Yibo Zhang, Weiwei Xu, Rui Ma 等CVPR 2025
- MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud RegistrationShuyuan Lin, Wenwu Peng, Junjie Huang, Qiang Qi 等AAAI 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt 等NeurIPS 2021 · 被引用 2,500 次
相关 Paper
- DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View DiffusionYansong Qu, Shaohui Dai, Xinyang Li, Yuze Wang 等AAAI 2026
- MGD: Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D VideoWeixing Xie, Ying Ye, Xian Wu, Jintian Li 等AAAI 2026
- RoHM: Robust Human Motion Reconstruction via DiffusionSiwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler 等CVPR 2024 · 被引用 11 次
- MorpheuS: Neural Dynamic Surface Reconstruction from Monocular RGB-D VideoHengyi Wang, Jingwen Wang, Lourdes AgapitoCVPR 2024 · 被引用 4 次
- OccFusion: Rendering Occluded Humans with Generative Diffusion PriorsAdam Sun, Tiange Xiang, Scott L. Delp, Li Fei-Fei 等NeurIPS 2024 · 被引用 13 次
