From Points to Multi-Object 3D Reconstruction
Francis Engelmann, Konstantinos Rematas, Bastian Leibe, Vittorio Ferrari
Abstract
We propose a method to detect and reconstruct multiple 3D objects from a single RGB image. The key idea is to optimize for detection, alignment and shape jointly over all objects in the RGB image, while focusing on realistic and physically plausible reconstructions. To this end, we propose a key-point detector that localizes objects as center points and directly predicts all object properties, including 9-DoF bounding boxes and 3D shapes – all in a single forward pass. The proposed method formulates 3D shape reconstruction as a shape selection problem, i.e. it selects among exemplar shapes from a given database. This makes it agnostic to shape representations, which enables a lightweight reconstruction of realistic and visually-pleasing shapes based on CAD-models, while the training objective is formulated around point clouds and voxel representations. A collision-loss promotes non-intersecting objects, further increasing the reconstruction realism. Given the RGB image, the presented approach performs lightweight reconstruction in a single-stage, it is real-time capable, fully differentiable and end-to-end trainable. Our experiments compare multiple approaches for 9-DoF bounding box estimation, evaluate the novel shape-selection mechanism and compare to recent methods in terms of 3D bounding box estimation and 3D shape reconstruction quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- ROCA: Robust CAD Model Retrieval and Alignment from a Single ImageCan Gümeli, Angela Dai, Matthias NießnerCVPR 2022 · 43 citations
- Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageFeng Liu, Xiaoming LiuNeurIPS 2021 · 43 citations
- MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationZhifei Yang, Keyang Lu, Chao Zhang, Jiaxing Qi et al.AAAI 2025 · 21 citations
- S-NeRF: Neural Radiance Fields for Street ViewsZiyang Xie, Junge Zhang, Wenye Li, Feihu Zhang et al.ICLR 2023 · 13 citations
- Multi-Object Manipulation via Object-Centric Neural Scattering FunctionsStephen Tian, Yancheng Cai, Hong-Xing Yu, Sergey Zakharov et al.CVPR 2023
Builds on10
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Escaping Plato's Cave: 3D Shape From Adversarial RenderingPhilipp Henzler, Niloy J. Mitra, Tobias RitschelICCV 2019 · 254 citations
- An Analysis of SVD for Deep Rotation EstimationJake Levinson, Carlos Esteves, Kefan Chen, Noah Snavely et al.NeurIPS 2020 · 131 citations
Related papers
- Learning Local RGB-to-CAD Correspondences for Object Pose EstimationGeorgios Georgakis, Srikrishna Karanam, Ziyan Wu, Jana KoseckaICCV 2019 · 25 citations
- Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single ImageWeicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela DaiICCV 2021 · 42 citations
- Single Image Shape-from-SilhouettesYawen Lu, Yuxing Wang, Guoyu LuACM MM 2020 · 6 citations
- Detection Based Part-level Articulated Object Reconstruction from Single RGBD ImageYuki Kawana, Tatsuya HaradaNeurIPS 2023 · 20 citations
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
