Unified Category-Level Object Detection and Pose Estimation from RGB Images Using 3D Prototypes
Tom Fischer, Xiaojie Zhang, Eddy Ilg
Abstract
Recognizing objects in images is a fundamental problem in computer vision. Although detecting objects in 2D images is common, many applications require determining their pose in 3D space. Traditional category-level methods rely on RGB-D inputs, which may not always be available, or employ two-stage approaches that use separate models and representations for detection and pose estimation. For the first time, we introduce a unified model that integrates detection and pose estimation into a single framework for RGB images by leveraging neural mesh models with learned features and multi-model RANSAC. Our approach achieves state-of-the-art results for RGB category-level pose estimation on REAL275, improving on the current state-of-the-art by 22.9% averaged across all scale-agnostic metrics. Finally, we demonstrate that our unified method exhibits greater robustness compared to single-stage baselines. Our code and models are available at https://github.com/Fischer-Tom/unified-detection-and-pose-estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b53ccd7-0852-45c3-9753-902dc4d511acBuilds on14
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose EstimationYongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach et al.CVPR 2022 · 170 citations
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu et al.ICCV 2021 · 169 citations
Related papers
- Mesh R-CNNGeorgia Gkioxari, Justin Johnson, Jitendra MalikICCV 2019
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
- DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldHaowen Wang, Zhipeng Fan, Zhen Zhao, Zhengping Che et al.ACM MM 2023 · 6 citations
- Unsupervised Learning of Category-Level 3D Pose from Object-Centric VideosLeonhard Sommer, Artur Jesslen, Eddy Ilg, Adam KortylewskiCVPR 2024
- Simultaneous Scene-independent Camera Localization and Category-level Object Pose Estimation via Multi-level Feature FusionJunyi Wang, Yue QiIEEE VR 2023 · 6 citations
