ODAM: Object Detection, Association, and Mapping using Posed RGB Video
Kejie Li, Daniel DeTone, Steven Chen, Minh Vo, Ian Reid, Hamid Rezatofighi, Chris Sweeney, Julian Straub, Richard A. Newcombe
Abstract
Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system relies on a deep learning front-end to detect 3D objects from a given RGB frame and associate them to a global object-based map using a graph neural network (GNN). Based on these frame-to-model associations, our back-end optimizes object bounding volumes, represented as super-quadrics, under multi-view geometry constraints and the object scale prior. We validate the proposed system on ScanNet where we show a significant improvement over existing RGB-only methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af9d8bb9-8ce8-4add-ad12-d555cdbdf46eCited by top-tier papers8
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao et al.NeurIPS 2025 · 26 citations
- Pixel-Aligned Recurrent Queries for Multi-View 3D Object DetectionYiming Xie, Huaizu Jiang, Georgia Gkioxari, Julian StraubICCV 2023 · 15 citations
- CAD-Estate: Large-scale CAD Model Annotation in RGB VideosKevis-Kokitsi Maninis, Stefan Popov, Matthias Nießner, Vittorio FerrariICCV 2023 · 14 citations
- Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D EnvironmentsLiyuan Zhu, Shengyu Huang, Konrad Schindler, Iro ArmeniCVPR 2024 · 10 citations
- GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking SystemShuo Wang, Yongcai Wang, Zhimin Xu, Yongyu Guo et al.ACM MM 2024 · 6 citations
Builds on8
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- ImVoteNet: Boosting 3D Object Detection in Point Clouds With Image VotesCharles R. Qi, Xinlei Chen, Or Litany, Leonidas J. GuibasCVPR 2020
Related papers
- FroDO: From Detections to 3D ObjectsMartin Rünz, Kejie Li, Meng Tang, Lingni Ma et al.CVPR 2020
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and MappingJustin Lazarow, Kai Kang, Afshin DehghanNeurIPS 2025 · 2 citations
- DSGN: Deep Stereo Geometry Network for 3D Object DetectionYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaCVPR 2020
- 3DP3: 3D Scene Perception via Probabilistic ProgrammingNishad Gothoskar, Marco F. Cusumano-Towner, Ben Zinberg, Matin Ghavamizadeh et al.NeurIPS 2021 · 59 citations
- ELLIPSDF: Joint Object Pose and Shape Optimization with a Bi-level Ellipsoid and Signed Distance Function DescriptionMo Shan, Qiaojun Feng, You-Yi Jau, Nikolay AtanasovICCV 2021 · 14 citations
