3DP3: 3D Scene Perception via Probabilistic Programming
Nishad Gothoskar, Marco F. Cusumano-Towner, Ben Zinberg, Matin Ghavamizadeh, Falk Pollok, Austin Garrett, Josh Tenenbaum, Dan Gutfreund, Vikash K. Mansinghka
Abstract
We present 3DP3, a framework for inverse graphics that uses inference in a structured generative model of objects, scenes, and images. 3DP3 uses (i) voxel models to represent the 3D shape of objects, (ii) hierarchical scene graphs to decompose scenes into objects and the contacts between them, and (iii) depth image likelihoods based on real-time graphics. Given an observed RGB-D image, 3DP3's inference algorithm infers the underlying latent 3D scene, including the object poses and a parsimonious joint parametrization of these poses, using fast bottom-up pose proposals, novel involutive MCMC updates of the scene graph structure, and, optionally, neural object detectors and pose estimators. We show that 3DP3 enables scene understanding that is aware of 3D shape, occlusion, and contact structure. Our results demonstrate that 3DP3 is more accurate at 6DoF object pose estimation from real images than deep learning baselines and shows better generalization to challenging scenes with novel viewpoints, contact, and partial observability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59883660-9f14-4e3c-973e-5b64ebb68070Cited by top-tier papers8
- VAEL: Bridging Variational Autoencoders and Probabilistic Logic ProgrammingEleonora Misino, Giuseppe Marra, Emanuele SansoneNeurIPS 2022 · 38 citations
- LILO: Learning Interpretable Libraries by Compressing and Documenting CodeGabriel Grand, Lionel Wong, Matthew Bowers, Theo X. Olausson et al.ICLR 2024 · 35 citations
- Designing Perceptual Puzzles by Differentiating Probabilistic ProgramsKartik Chandra, Tzu-Mao Li, Joshua B. Tenenbaum, Jonathan Ragan-KelleySIGGRAPH 2022 · 17 citations
- Probabilistic Programming with Stochastic ProbabilitiesAlexander K. Lew, Matin Ghavamizadeh, Martin C. Rinard, Vikash K. MansinghkaPLDI 2023 · 9 citations
- Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-imageYu Zhao, Hao Fei, Xiangtai Li, Libo Qin et al.NeurIPS 2024 · 2 citations
Builds on4
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu et al.ICCV 2019 · 130 citations
- 3D-RelNet: Joint Object and Relational Network for 3D PredictionNilesh Kulkarni, Ishan Misra, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 48 citations
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 10 citations
Related papers
- 3D Neural Embedding Likelihood: Probabilistic Inverse Graphics for Robust 6D Pose EstimationGuangyao Zhou, Nishad Gothoskar, Lirui Wang, Joshua B. Tenenbaum et al.ICCV 2023 · 5 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingShaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi et al.CVPR 2020
- EVA3D: Compositional 3D Human Generation from 2D Image CollectionsFangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan et al.ICLR 2023 · 35 citations
- MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric FusionKentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton et al.CVPR 2020
