Open-vocabulary object 6D pose estimation
Jaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro, Fabio Poiesi
摘要
We introduce the new setting of open-vocabulary object 6D pose estimation, in which a textual prompt is used to specify the object of interest. In contrast to existing approaches, in our setting (i) the object of interest is specified solely through the textual prompt, (ii) no object model (e.g., CAD or video sequence) is required at inference, and (iii) the object is imaged from two RGBD viewpoints of different scenes. To operate in this setting, we introduce a novel approach that leverages a Vision-Language Model to segment the object of interest from the scenes and to estimate its relative 6D pose. The key of our approach is a carefully devised strategy to fuse object-level information provided by the prompt with local image features, resulting in a feature space that can generalize to novel concepts. We validate our approach on a new benchmark based on two popular datasets, REAL275 and Toyota-Light, which collectively encompass 34 object instances appearing in four thousand image pairs. The results demonstrate that our approach outperforms both a well-established handcrafted method and a recent deep learning-based baseline in estimating the relative 6D pose of objects in different scenes. Code and dataset are available at https: //jcorsetti.github.io/oryon .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept VectorsLiming Kuang, Yordanka Velikova, Mahdi Saleh, Jan-Nico Zaech 等CVPR 2026 · 被引用 5 次
- SingRef6D: Monocular Novel Object Pose Estimation with a Single RGB ReferenceJiahui Wang, Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 等NeurIPS 2025 · 被引用 2 次
- Zero-Shot Inexact CAD Model Alignment from a Single ImagePattaramanee Arsomngern, Sasikarn Khwanmuang, Matthias Nießner, Supasorn SuwajanakornICCV 2025 · 被引用 2 次
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map GenerationDexin Zuo, Ang Li, Wei Wang, Wenxian Yu 等AAAI 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Fully Convolutional Geometric FeaturesChristopher B. Choy, Jaesik Park, Vladlen KoltunICCV 2019 · 被引用 807 次
- REGTR: End-to-end Point Cloud Correspondences with TransformersZi Jian Yew, Gim Hee LeeCVPR 2022 · 被引用 242 次
- OMNet: Learning Overlapping Mask for Partial-to-Partial Point Cloud RegistrationHao Xu, Shuaicheng Liu, Guangfu Wang, Guanghui Liu 等ICCV 2021 · 被引用 195 次
相关 Paper
- CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific RepresentationsHua Wang, Hong Liu, Jiale Ren, Mingxin Tan 等ACM MM 2025 · 被引用 1 次
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 被引用 215 次
- Vision Foundation Model Enables Generalizable Object Pose EstimationKai Chen, Yiyao Ma, Xingyu Lin, Stephen James 等NeurIPS 2024 · 被引用 5 次
- Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship DetectionShuo Yang, Yongqi Wang, Xiaofeng Ji, Xinxiao WuAAAI 2024 · 被引用 4 次
- Learning Deep Network for Detecting 3D Object Keypoints and 6D PosesWanqing Zhao, Shaobo Zhang, Ziyu Guan, Wei Zhao 等CVPR 2020
