FroDO: From Detections to 3D Objects
Martin Rünz, Kejie Li, Meng Tang, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
Abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a50f49d-2d1b-458d-8278-1e4c488716ceCited by top-tier papers18
- CodeNeRF: Disentangled Neural Radiance Fields for Object CategoriesWonbong Jang, Lourdes AgapitoICCV 2021 · 246 citations
- Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageFeng Liu, Xiaoming LiuNeurIPS 2021 · 43 citations
- PINs: Progressive Implicit Networks for Multi-Scale Neural RepresentationsZoe Landgraf, Alexander Sorkine-Hornung, Ricardo Silveira CabralICML 2022 · 24 citations
- Pixel-Aligned Recurrent Queries for Multi-View 3D Object DetectionYiming Xie, Huaizu Jiang, Georgia Gkioxari, Julian StraubICCV 2023 · 15 citations
- ELLIPSDF: Joint Object Pose and Shape Optimization with a Bi-level Ellipsoid and Signed Distance Function DescriptionMo Shan, Qiaojun Feng, You-Yi Jau, Nikolay AtanasovICCV 2021 · 14 citations
Builds on2
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingShaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi et al.CVPR 2020
Related papers
- CARTO: Category and Joint Agnostic Reconstruction of ARTiculated ObjectsNick Heppert, Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu et al.CVPR 2023
- ODAM: Object Detection, Association, and Mapping using Posed RGB VideoKejie Li, Daniel DeTone, Steven Chen, Minh Vo et al.ICCV 2021 · 31 citations
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and MappingJustin Lazarow, Kai Kang, Afshin DehghanNeurIPS 2025 · 2 citations
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB VideosZiren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han et al.CVPR 2026 · 13 citations
- RfD-Net: Point Scene Understanding by Semantic Instance ReconstructionYinyu Nie, Ji Hou, Xiaoguang Han, Matthias NießnerCVPR 2021
