CLIP-AdaM: Adapting Multi-view CLIP for Open-set 3D Object Retrieval
Xinwei He, Liang Ma, Yuxuan Cheng, Zhichuan Wang, Yulong Wang, Yang Zhou, Xiang Bai
Abstract
Open-set 3D object retrieval (3DOR) aims to learn discriminative and generalizable embeddings for unseen categories of 3D objects. However, attaining this objective typically requires the costly acquisition of large-scale 3D object datasets and associated resources for model training. Building upon the strong open-world representation capabilities of CLIP, we introduce CLIP-AdaM, which, to our knowledge, represents the first attempt to adapt a CLIP model for open-set 3DOR with minimal effort. We first find that a pretrained CLIP already delivers a surprisingly acceptable performance on multi-view images. To further unleash its potential, we design a customized adapter for learning to aggregate and adapt its pretrained features towards better 3D embeddings. For aggregation, it learns two sets of view scores to weigh the contributions of view images for fusion. One is learned by a tiny view-score network at the instance level, and the other is learned implicitly at the dataset level, aiding generalization to unseen categories. The adaptation component comprises only a basic linear layer yet yields superior results. During training, the adapter with such a small amount of parameters can be efficiently fine-tuned with limited 3D closed-set data, effectively mitigating the overfitting issue while harnessing the prior knowledge from pretrained models. Without bells and whistles, CLIP-AdaM attains state-of-the-art performance on four open-set 3DOR benchmarks. Additionally, it demonstrates strong extensibility to broader scenarios, including zero-shot, few-shot, and seen/unseen 3D representation learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object RetrievalZhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu et al.ICCV 2025 · 2 citations
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo et al.ICCV 2023 · 248 citations
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
- CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific RepresentationsHua Wang, Hong Liu, Jiale Ren, Mingxin Tan et al.ACM MM 2025 · 1 citation
- Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCheng Cheng, Lin Song, Ruoyi Xue, Hang Wang et al.NeurIPS 2023 · 65 citations
