Vision Foundation Model Enables Generalizable Object Pose Estimation
Kai Chen, Yiyao Ma, Xingyu Lin, Stephen James, Jianshu Zhou, Yun-Hui Liu, Pieter Abbeel, Qi Dou
摘要
Object pose estimation plays a crucial role in robotic manipulation, however, its practical applicability still suffers from limited generalizability. This paper addresses the challenge of generalizable object pose estimation, particularly focusing on category-level object pose estimation for unseen object categories. Current meth-ods either require impractical instance-level training or are confined to predefined categories, limiting their applicability. We propose VFM-6D, a novel framework that explores harnessing existing vision and language models, to elaborate object pose estimation into two stages: category-level object viewpoint estimation and object coordinate map estimation. Based on the two-stage framework, we introduce a 2D-to-3D feature lifting module and a shape-matching module, both of which leverage pre-trained vision foundation models to improve object representation and matching accuracy. VFM-6D is trained on cost-effective synthetic data and exhibits superior generalization capabilities. It can be applied to both instance-level unseen object pose estimation and category-level object pose estimation for novel categories. Evaluations on benchmark datasets demonstrate the effectiveness and versatility of VFM-6D in various real-world scenarios. Project website: vfm-6d.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose EstimationYuchen Che, JINGTU WU, Hao ZHENG, Asako KanezakiCVPR 2026 · 被引用 1 次
- Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation LearningYIYAO MA, Kai Chen, Zhongxiang Zhou, Zhuheng Song 等ICML 2026
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 被引用 215 次
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept VectorsLiming Kuang, Yordanka Velikova, Mahdi Saleh, Jan-Nico Zaech 等CVPR 2026 · 被引用 5 次
- Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose EstimationXiao Lin, Wenfei Yang, Yuan Gao, Tianzhu ZhangCVPR 2024
- Universal Features Guided Zero-Shot Category-Level Object Pose EstimationWentian Qu, Chenyu Meng, Heng Li, Jian Cheng 等AAAI 2025
- Self-Supervised Category-Level 6D Object Pose Estimation with Deep Implicit Shape RepresentationWanli Peng, Jianhang Yan, Hongtao Wen, Yi SunAAAI 2022 · 被引用 46 次
