FroDO: From Detections to 3D Objects
Martin Rünz, Kejie Li, Meng Tang, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
摘要
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- CodeNeRF: Disentangled Neural Radiance Fields for Object CategoriesWonbong Jang, Lourdes AgapitoICCV 2021 · 被引用 246 次
- Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageFeng Liu, Xiaoming LiuNeurIPS 2021 · 被引用 43 次
- PINs: Progressive Implicit Networks for Multi-Scale Neural RepresentationsZoe Landgraf, Alexander Sorkine-Hornung, Ricardo Silveira CabralICML 2022 · 被引用 24 次
- Pixel-Aligned Recurrent Queries for Multi-View 3D Object DetectionYiming Xie, Huaizu Jiang, Georgia Gkioxari, Julian StraubICCV 2023 · 被引用 15 次
- ELLIPSDF: Joint Object Pose and Shape Optimization with a Bi-level Ellipsoid and Signed Distance Function DescriptionMo Shan, Qiaojun Feng, You-Yi Jau, Nikolay AtanasovICCV 2021 · 被引用 14 次
它引用的顶会 Paper2
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou 等ICCV 2019 · 被引用 373 次
- DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingShaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi 等CVPR 2020
相关 Paper
- CARTO: Category and Joint Agnostic Reconstruction of ARTiculated ObjectsNick Heppert, Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu 等CVPR 2023
- ODAM: Object Detection, Association, and Mapping using Posed RGB VideoKejie Li, Daniel DeTone, Steven Chen, Minh Vo 等ICCV 2021 · 被引用 31 次
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and MappingJustin Lazarow, Kai Kang, Afshin DehghanNeurIPS 2025 · 被引用 2 次
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB VideosZiren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han 等CVPR 2026 · 被引用 13 次
- RfD-Net: Point Scene Understanding by Semantic Instance ReconstructionYinyu Nie, Ji Hou, Xiaoguang Han, Matthias NießnerCVPR 2021
