O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model
Yubin Hu, Sheng Ye, Wang Zhao, Matthieu Lin, Yuze He, Yu-Hui Wen, Ying He, Yong-Jin Liu
Abstract
Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc731022-9fda-49aa-8c74-b053d653b23fCited by top-tier papers4
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction From Cluttered ScenesZesong Yang, Bangbang Yang, Liyuan Cui, Yuewen Ma et al.ICCV 2025 · 2 citations
- MIDGArD: Modular Interpretable Diffusion over Graphs for Articulated DesignsQuentin Leboutet, Nina Wiedemann, Zhipeng Cai, Michael Paulitsch et al.NeurIPS 2024 · 2 citations
- DecoupledGaussian: Object-Scene Decoupling for Physics-Based InteractionMiaowei Wang, Yibo Zhang, Weiwei Xu, Rui Ma et al.CVPR 2025
- MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud RegistrationShuyuan Lin, Wenwu Peng, Junjie Huang, Qiang Qi et al.AAAI 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
Related papers
- DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View DiffusionYansong Qu, Shaohui Dai, Xinyang Li, Yuze Wang et al.AAAI 2026
- MGD: Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D VideoWeixing Xie, Ying Ye, Xian Wu, Jintian Li et al.AAAI 2026
- RoHM: Robust Human Motion Reconstruction via DiffusionSiwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler et al.CVPR 2024 · 11 citations
- MorpheuS: Neural Dynamic Surface Reconstruction from Monocular RGB-D VideoHengyi Wang, Jingwen Wang, Lourdes AgapitoCVPR 2024 · 4 citations
- OccFusion: Rendering Occluded Humans with Generative Diffusion PriorsAdam Sun, Tiange Xiang, Scott L. Delp, Li Fei-Fei et al.NeurIPS 2024 · 13 citations
