DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
Xinwei He, Yansong Zheng, Qianru Han, Zhichuan Wang, Yuxuan Cai, Yang Zhou, Jingbo Xia, Yulong Wang, Jinhai Xiang, Xiang Bai
摘要
Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to build view-based 3D descriptors. Despite CLIP's strong generalization ability, its lack of fine-grainedness prompted us to explore the potential of a more recent self-supervised encoder-DINO. To address this, we propose DINO Eats CLIP (DEC), a novel framework for dynamic multi-view integration that is regularized by synthesizing data for unseen classes. We first find that simply mean-pooling over view features from a frozen DINO backbone gives decent performance. Yet, further adaptation causes severe overfitting on average view patterns of known classes. To combat it, we then design a module named Chunking and Adapting Module (CAM). It segments multi-view images into chunks and dynamically integrates local view relations, yielding more robust features than the standard pooling strategy. Finally, we propose Virtual Feature Synthesis (VFS) module to mitigate bias towards known categories explicitly. Under the hood, VFS leverages CLIP's broad, pre-aligned vision-language space to synthesize virtual features for unseen classes. By exposing DEC to these virtual features, we greatly enhance its open-set discrimination capacity. Extensive experiments on standard open-set 3DOR benchmarks demonstrate its superior efficacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- VOS: Learning What You Don't Know by Virtual Outlier SynthesisXuefeng Du, Zhaoning Wang, Mu Cai, Yixuan LiICLR 2022 · 被引用 417 次
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu 等NeurIPS 2023 · 被引用 267 次
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo 等ICCV 2023 · 被引用 248 次
相关 Paper
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object RetrievalZhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu 等ICCV 2025 · 被引用 2 次
- CLIP-AdaM: Adapting Multi-view CLIP for Open-set 3D Object RetrievalXinwei He, Liang Ma, Yuxuan Cheng, Zhichuan Wang 等SIGIR 2025 · 被引用 3 次
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionJunjie Wang, Bin Chen, Yulin Li, Bin Kang 等CVPR 2025
- Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsHeeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho 等NeurIPS 2024 · 被引用 32 次
- DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language AlignmentCijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre 等CVPR 2025
