Peek-a-Boo: Occlusion Reasoning in Indoor Scenes With Plane Representations
Ziyu Jiang, Buyu Liu, Samuel Schulter, Zhangyang Wang, Manmohan Chandraker
Abstract
We address the challenging task of occlusion-aware indoor 3D scene understanding. We represent scenes by a set of planes, where each one is defined by its normal, offset and two masks outlining (i) the extent of the visible part and (ii) the full region that consists of both visible and occluded parts of the plane. We infer these planes from a single input image with a novel neural network architecture. It consists of a twobranch category-specific module that aims to predict layout and objects of the scene separately so that different types of planes can be handled better. We also introduce a novel loss function based on plane warping that can leverage multiple views at training time for improved occlusion-aware reasoning. In order to train and evaluate our occlusion-reasoning model, we use the ScanNet dataset [1] and propose (i) a strategy to automatically extract ground truth for both visible and hidden regions and (ii) a new evaluation metric that specifically focuses on the prediction in hidden regions. We empirically demonstrate that our proposed approach can achieve higher accuracy for occlusion reasoning compared to competitive baselines on the ScanNet dataset, e.g. 42.65% relative improvement on hidden regions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9413d361-c886-49d0-8fe7-2f8b1b025a67Cited by top-tier papers8
- PlaneTR: Structure-Guided Transformers for 3D Plane RecoveryBin Tan, Nan Xue, Song Bai, Tianfu Wu et al.ICCV 2021 · 51 citations
- Planar Surface Reconstruction from Sparse ViewsLinyi Jin, Shengyi Qian, Andrew Owens, David F. FouheyICCV 2021 · 51 citations
- ESCNet: Gaze Target Detection with the Understanding of 3D ScenesJun Bao, Buyu Liu, Jun YuCVPR 2022 · 36 citations
- Understanding 3D Object Articulation in Internet VideosShengyi Qian, Linyi Jin, Chris Rockwell, Siyi Chen et al.CVPR 2022 · 13 citations
- Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionBowen Zhang, Qing Liu, Jianming Zhang, Yilin Wang et al.AAAI 2024 · 4 citations
Builds on1
Related papers
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image GenerationVaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla et al.CVPR 2026 · 5 citations
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- PlaneRAS: Learning Planar Primitives for 3D Plane RecoveryFang Zhang, Wenzhao Zheng, Linqing Zhao, Zelan Zhu et al.ICCV 2025
- SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language PretrainingYue Li, Qi Ma, Runyi Yang, Huapeng Li et al.ICCV 2025 · 5 citations
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng et al.CVPR 2021
