3DP3: 3D Scene Perception via Probabilistic Programming
Nishad Gothoskar, Marco F. Cusumano-Towner, Ben Zinberg, Matin Ghavamizadeh, Falk Pollok, Austin Garrett, Josh Tenenbaum, Dan Gutfreund, Vikash K. Mansinghka
摘要
We present 3DP3, a framework for inverse graphics that uses inference in a structured generative model of objects, scenes, and images. 3DP3 uses (i) voxel models to represent the 3D shape of objects, (ii) hierarchical scene graphs to decompose scenes into objects and the contacts between them, and (iii) depth image likelihoods based on real-time graphics. Given an observed RGB-D image, 3DP3's inference algorithm infers the underlying latent 3D scene, including the object poses and a parsimonious joint parametrization of these poses, using fast bottom-up pose proposals, novel involutive MCMC updates of the scene graph structure, and, optionally, neural object detectors and pose estimators. We show that 3DP3 enables scene understanding that is aware of 3D shape, occlusion, and contact structure. Our results demonstrate that 3DP3 is more accurate at 6DoF object pose estimation from real images than deep learning baselines and shows better generalization to challenging scenes with novel viewpoints, contact, and partial observability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VAEL: Bridging Variational Autoencoders and Probabilistic Logic ProgrammingEleonora Misino, Giuseppe Marra, Emanuele SansoneNeurIPS 2022 · 被引用 38 次
- LILO: Learning Interpretable Libraries by Compressing and Documenting CodeGabriel Grand, Lionel Wong, Matthew Bowers, Theo X. Olausson 等ICLR 2024 · 被引用 35 次
- Designing Perceptual Puzzles by Differentiating Probabilistic ProgramsKartik Chandra, Tzu-Mao Li, Joshua B. Tenenbaum, Jonathan Ragan-KelleySIGGRAPH 2022 · 被引用 17 次
- Probabilistic Programming with Stochastic ProbabilitiesAlexander K. Lew, Matin Ghavamizadeh, Martin C. Rinard, Vikash K. MansinghkaPLDI 2023 · 被引用 9 次
- Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-imageYu Zhao, Hao Fei, Xiangtai Li, Libo Qin 等NeurIPS 2024 · 被引用 2 次
它引用的顶会 Paper4
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak 等ICCV 2019 · 被引用 474 次
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu 等ICCV 2019 · 被引用 130 次
- 3D-RelNet: Joint Object and Relational Network for 3D PredictionNilesh Kulkarni, Ishan Misra, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 被引用 48 次
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 被引用 10 次
相关 Paper
- 3D Neural Embedding Likelihood: Probabilistic Inverse Graphics for Robust 6D Pose EstimationGuangyao Zhou, Nishad Gothoskar, Lirui Wang, Joshua B. Tenenbaum 等ICCV 2023 · 被引用 5 次
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 被引用 486 次
- DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingShaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi 等CVPR 2020
- EVA3D: Compositional 3D Human Generation from 2D Image CollectionsFangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan 等ICLR 2023 · 被引用 35 次
- MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric FusionKentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton 等CVPR 2020
