Unified Category-Level Object Detection and Pose Estimation from RGB Images Using 3D Prototypes
Tom Fischer, Xiaojie Zhang, Eddy Ilg
摘要
Recognizing objects in images is a fundamental problem in computer vision. Although detecting objects in 2D images is common, many applications require determining their pose in 3D space. Traditional category-level methods rely on RGB-D inputs, which may not always be available, or employ two-stage approaches that use separate models and representations for detection and pose estimation. For the first time, we introduce a unified model that integrates detection and pose estimation into a single framework for RGB images by leveraging neural mesh models with learned features and multi-model RANSAC. Our approach achieves state-of-the-art results for RGB category-level pose estimation on REAL275, improving on the current state-of-the-art by 22.9% averaged across all scale-agnostic metrics. Finally, we demonstrate that our unified method exhibits greater robustness compared to single-stage baselines. Our code and models are available at https://github.com/Fischer-Tom/unified-detection-and-pose-estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 被引用 486 次
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 被引用 183 次
- ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose EstimationYongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach 等CVPR 2022 · 被引用 170 次
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu 等ICCV 2021 · 被引用 169 次
相关 Paper
- Mesh R-CNNGeorgia Gkioxari, Justin Johnson, Jitendra MalikICCV 2019
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
- DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldHaowen Wang, Zhipeng Fan, Zhen Zhao, Zhengping Che 等ACM MM 2023 · 被引用 6 次
- Unsupervised Learning of Category-Level 3D Pose from Object-Centric VideosLeonhard Sommer, Artur Jesslen, Eddy Ilg, Adam KortylewskiCVPR 2024
- Simultaneous Scene-independent Camera Localization and Category-level Object Pose Estimation via Multi-level Feature FusionJunyi Wang, Yue QiIEEE VR 2023 · 被引用 6 次
