Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes From a Single Image
Yinyu Nie, Xiaoguang Han, Shihui Guo, Yujian Zheng, Jian Chang, Jian-Jun Zhang
Abstract
Semantic reconstruction of indoor scenes refers to both scene understanding and object reconstruction. Existing works either address one part of this problem or focus on independent objects. In this paper, we bridge the gap between understanding and reconstruction, and propose an end-to-end solution to jointly reconstruct room layout, object bounding boxes and meshes from a single image. Instead of separately resolving scene understanding and object reconstruction, our method builds upon a holistic scene context and proposes a coarse-to-fine hierarchy with three components: 1. room layout with camera pose; 2. 3D object bounding boxes; 3. object meshes. We argue that understanding the context of each component can assist the task of parsing the others, which enables joint understanding and reconstruction. The experiments on the SUN RGB-D and Pix3D datasets demonstrate that our method consistently outperforms existing methods in indoor layout estimation, 3D object detection and mesh reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f14fef3-6f2f-4e26-ac21-721c4c821a7dCited by top-tier papers97
- Neural RGB-D Surface ReconstructionDejan Azinovic, Ricardo Martin-Brualla, Dan B. Goldman, Matthias Nießner et al.CVPR 2022 · 272 citations
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- Panoptic Neural Fields: A Semantic Object-Aware Neural Scene RepresentationAbhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi et al.CVPR 2022 · 204 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree NetworksZhihao Liang, Zhihao Li, Songcen Xu, Mingkui Tan et al.ICCV 2021 · 170 citations
Builds on5
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu et al.ICCV 2019 · 130 citations
- 3D Scene Reconstruction With Multi-Layer Depth and Epipolar TransformersDaeyun Shin, Zhile Ren, Erik B. Sudderth, Charless C. FowlkesICCV 2019 · 67 citations
- Few-Shot Generalization for Single-Image 3D Reconstruction via PriorsBram Wallace, Bharath HariharanICCV 2019 · 43 citations
- Mesh R-CNNGeorgia Gkioxari, Justin Johnson, Jitendra MalikICCV 2019
Related papers
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
- PanoContext-Former: Panoramic Total Scene Understanding with a TransformerYuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong et al.CVPR 2024
- PixARMesh: Autoregressive Mesh-Native Single-View Scene ReconstructionXiang Zhang, Sohyun Yoo, Hongrui Wu, Chuan Li et al.CVPR 2026 · 2 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
