CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention
Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma, Xupeng Miao, Xuming He, Bin Cui
Abstract
Contrastive Language-Image Pre-training (CLIP) has been shown to learn visual representations with promising zero-shot performance. To further improve its downstream accuracy, existing works propose additional learnable modules upon CLIP and fine-tune them by few-shot training sets. However, the resulting extra training cost and data requirement severely hinder the efficiency for model deployment and knowledge transfer. In this paper, we introduce a free-lunch enhancement method, CALIP, to boost CLIP's zero-shot performance via a parameter-free attention module. Specifically, we guide visual and textual representations to interact with each other and explore cross-modal informative features via attention. As the pre-training has largely reduced the embedding distances between two modalities, we discard all learnable parameters in the attention and bidirectionally update the multi-modal features, enabling the whole process to be parameter-free and training-free. In this way, the images are blended with textual-aware signals and the text representations become visual-guided for better adaptive zero-shot alignment. We evaluate CALIP on various benchmarks of 14 datasets for both 2D image and 3D point cloud few-shot classification, showing consistent zero-shot performance improvement over CLIP. Based on that, we further insert a small number of linear layers in CALIP's attention module and verify our robustness under the few-shot settings, which also achieves leading performance compared to existing methods. Those extensive experiments demonstrate the superiority of our approach for efficient enhancement of CLIP. Code is available at https://github.com/ZiyuGuo99/CALIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03e8dc65-a0dc-4b10-a5b3-f53cab1bf480Cited by top-tier papers54
- SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsVishaal Udandarao, Ankush Gupta, Samuel AlbanieICCV 2023 · 160 citations
- Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementXiangyang Zhu, Renrui Zhang, Bowei He, Aojun Zhou et al.ICCV 2023 · 121 citations
- Prompt-Based Distribution Alignment for Unsupervised Domain AdaptationShuanghao Bai, Min Zhang, Wanqi Zhou, Siteng Huang et al.AAAI 2024 · 103 citations
- Enhance Vision-Language Alignment with NoiseSida Huang, Hongyuan Zhang, Xuelong LiAAAI 2025 · 96 citations
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang et al.ICCV 2023 · 86 citations
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
Related papers
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
- Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIPYayuan Li, Jintao Guo, Lei Qi, Wenbin Li et al.AAAI 2025 · 9 citations
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- PointCLIP: Point Cloud Understanding by CLIPRenrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li et al.CVPR 2022
- CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingTianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang et al.ICCV 2023 · 220 citations
