Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP
Yayuan Li, Jintao Guo, Lei Qi, Wenbin Li, Yinghuan Shi
Abstract
Contrastive Language-Image Pretraining (CLIP) has been widely used in vision tasks. Notably, CLIP has demonstrated promising performance in few-shot learning (FSL). However, existing CLIP-based methods in training-free FSL (i.e., without the requirement of additional training) mainly learn different modalities independently, leading to two essential issues: 1) severe anomalous match in image modality; 2) varying quality of generated text prompts. To address these issues, we build a mutual guidance mechanism, that introduces an Image-Guided-Text (IGT) component to rectify varying quality of text prompts through image representations, and a Text-Guided-Image (TGI) component to mitigate the anomalous match of image modality through text representations. By integrating IGT and TGI, we adopt a perspective of Text-Image Mutual guidance Optimization, proposing TIMO. Extensive experiments show that TIMO significantly outperforms the state-of-the-art (SOTA) training-free method. Additionally, by exploring the extent of mutual guidance, we propose an enhanced variant, TIMO-S, which even surpasses the best training-required methods by 0.33% with approximately ×100 less time cost. Our code is available at https://github.com/lyymuwu/TIMO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dc24b50-1ae3-4d26-8339-1e983c18904cCited by top-tier papers4
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- When Shared Knowledge Hurts: Spectral Over-Accumulation in Model MergingYayuan Li, Ze Peng, Jian Zhang, Jintao Guo et al.ICML 2026 · 5 citations
- Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI DecodingShumeng Li, Jintao Guo, Jian Zhang, Yulin Zhou et al.CVPR 2026 · 2 citations
- DO: A Dual Debiasing Operator for Training-Free Test-Time Adaptation of Vision–Language ModelsYihong Luo, Wenwu He, Dong Liang, Yihang Zhou et al.ICML 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- What does a platypus look like? Generating customized prompts for zero-shot image classificationSarah M. Pratt, Ian Covert, Rosanne Liu, Ali FarhadiICCV 2023 · 343 citations
- Prompt Distribution LearningYuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu et al.CVPR 2022 · 212 citations
- SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsVishaal Udandarao, Ankush Gupta, Samuel AlbanieICCV 2023 · 160 citations
- Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsKarsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals et al.ICCV 2023 · 124 citations
Related papers
- CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionZiyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma et al.AAAI 2023 · 182 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen et al.CVPR 2024 · 27 citations
- Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts LearningQiang Wang, Junlong Du, Ke Yan, Shouhong DingACM MM 2023 · 26 citations
- Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Zhimeng Huang et al.AAAI 2026 · 3 citations
