RevealNet: Seeing Behind Objects in RGB-D Scans
Ji Hou, Angela Dai, Matthias Nießner
Abstract
During 3D reconstruction, it is often the case that people cannot scan each individual object from all views, resulting in missing geometry in the captured scan. This missing geometry can be fundamentally limiting for many applications, e.g., a robot needs to know the unseen geometry to perform a precise grasp on an object. Thus, we introduce the task of semantic instance completion: from an incomplete RGB-D scan of a scene, we aim to detect the individual object instances and infer their complete object geometry. This will open up new possibilities for interactions with objects in a scene, for instance for virtual or robotic agents. We tackle this problem by introducing RevealNet, a new data-driven approach that jointly detects object instances and predicts their complete geometry. This enables a semantically meaningful decomposition of a scanned scene into individual, complete 3D objects, including hidden and unobserved object parts. RevealNet is an end-to-end 3D neural network architecture that leverages joint color and geometry feature learning. The fully-convolutional nature of our 3D network enables efficient inference of semantic instance completion for 3D scans at scale of large indoor environments in a single forward pass. We show that predicting complete object geometry improves both 3D detection and instance segmentation performance. We evaluate on both real and synthetic scan benchmark data for the new task, where we outperform state-of-the-art approaches by over 15 in mAP@0.5 on ScanNet, and over 18 in mAP@0.5 on SUNCG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa34f5fc-5d9a-4fb7-8a4e-6b76ad75f877Cited by top-tier papers28
- Unsupervised Point Cloud Pre-training via Occlusion CompletionHanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby et al.ICCV 2021 · 323 citations
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai et al.ICCV 2021 · 94 citations
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner et al.ICCV 2023 · 82 citations
Builds on1
Related papers
- Towards Part-Based Understanding of RGB-D ScansAlexey Bokhovkin, Vladislav Ishimtsev, Emil Bogomolov, Denis Zorin et al.CVPR 2021
- SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D ScansAngela Dai, Christian Diller, Matthias NießnerCVPR 2020
- Point Cloud Semantic Scene Completion from RGB-D ImagesShoulong Zhang, Shuai Li, Aimin Hao, Hong QinAAAI 2021 · 13 citations
- Point-based Instance Completion with Scene ConstraintsWesley Khademi, Fuxin LiICLR 2025
- EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene SupervisionJiahao Chen, Zihui Zhang, Yafei Yang, Jinxi Li et al.CVPR 2026 · 1 citation
