Learning 3D Scene Priors with 2D Supervision
Yinyu Nie, Angela Dai, Xiaoguang Han, Matthias Nießner
Abstract
Holistic 3D scene understanding entails estimation of both layout configuration and object geometry in a 3D environment. Recent works have shown advances in 3D scene estimation from various input modalities (e.g., images, 3D scans), by leveraging 3D supervision (e.g., 3D bounding boxes or CAD models), for which collection at scale is expensive and often intractable. To address this shortcoming, we propose a new method to learn 3D scene priors of layout and shape without requiring any 3D ground truth. Instead, we rely on 2D supervision from multi-view RGB images. Our method represents a 3D scene as a latent vector, from which we can progressively decode to a sequence of objects characterized by their class categories, 3D bounding boxes, and meshes. With our trained autoregressive decoder representing the scene prior, our method facilitates many downstream applications, including scene synthesis, interpolation, and single-view reconstruction. Experiments on 3D-FRONT and ScanNet show that our method outperforms state of the art in single-view reconstruction, and achieves state-of-the-art results in scene synthesis against baselines which require for 3D supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbed5113-0373-43e1-9916-55577d571ea3Cited by top-tier papers15
- CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene GraphsGuangyao Zhai, Evin Pinar Örnek, Shun-Cheng Wu, Yan Di et al.NeurIPS 2023 · 76 citations
- DiffuScene: Denoising Diffusion Models for Generative Indoor Scene SynthesisJiapeng Tang, Yinyu Nie, Lev Markhasin, Angela Dai et al.CVPR 2024 · 62 citations
- NAP: Neural 3D Articulated Object PriorJiahui Lei, Congyue Deng, William B. Shen, Leonidas J. Guibas et al.NeurIPS 2023 · 53 citations
- Zero-Shot Scene Reconstruction from Single Images with Deep Prior AssemblyJunsheng Zhou, Yu-Shen Liu, Zhizhong HanNeurIPS 2024 · 42 citations
- PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AIYandan Yang, Baoxiong Jia, Peiyuan Zhi, Siyuan HuangCVPR 2024 · 27 citations
Builds on14
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsHuan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang et al.ICCV 2021 · 419 citations
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- ROCA: Robust CAD Model Retrieval and Alignment from a Single ImageCan Gümeli, Angela Dai, Matthias NießnerCVPR 2022 · 43 citations
Related papers
- Learning 3D Object Shape and Layout without 3D SupervisionGeorgia Gkioxari, Nikhila Ravi, Justin JohnsonCVPR 2022 · 18 citations
- Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single ImageWeicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela DaiICCV 2021 · 42 citations
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva et al.ICCV 2025 · 4 citations
- Discovering 3D Parts from Image CollectionsChun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan YangICCV 2021 · 21 citations
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
