CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific Representations
Hua Wang, Hong Liu, Jiale Ren, Mingxin Tan, Zhongzien Jiang
摘要
The open-vocabulary paradigm enhances 6D object pose estimation by leveraging language cues to transfer learned representations from seen to unseen objects, yet its performance suffers from a misalignment between vision-language representations and the 6D pose space. This challenge is compounded by the intrinsic lack of pose awareness in models like CLIP (Contrastive Language-Image Pre-Training). To address this limitation, we introduce CLIP-6D, which equips CLIP with the ability to estimate poses through the learning of generalizable representations specific to objects. CLIP-6D consists of three key components: (1) an innovative self-alignment strategy that enables CLIP to derive geometric representations from RGB images by leveraging its inherent feature extractor; % utilizing its inherent feature extraction capability; (2) a multiplex representation interactive learning method efficiently bridges the heterogeneous representations of object category priors, geometry, and spatial correlations; (3) a lightweight adapter using knowledge distillation improves CLIP's capture of detailed semantic representations. Experiments show that CLIP-6D achieves an improvement of 10.8% and 16.4% in the metric 5° 2 cm and 10° 5 cm for zero-shot generalization and achieves a speedup of 5.4 FPS in inference over current state-of-the-art methods. The source code and models are available at https://github.com/whoawong/CLIP-6D.git.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- ProposalCLIP: Unsupervised Open-Category Object Proposal Generation via Exploiting CLIP CuesHengcan Shi, Munawar Hayat, Yicheng Wu, Jianfei CaiCVPR 2022 · 被引用 59 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICLR 2024 · 被引用 129 次
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsShan Ning, Longtian Qiu, Yongfei Liu, Xuming HeCVPR 2023
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseXu Zhang, Wen Wang, Zhe Chen, Yufei Xu 等CVPR 2023
