Transductive Zero-Shot and Few-Shot CLIP
Ségolène Martin, Yunshi Huang, Fereshteh Shakeri, Jean-Christophe Pesquet, Ismail Ben Ayed
Abstract
Transductive inference has been widely investigated in few-shot image classification, but completely overlooked in the recent, fast growing literature on adapting vision-langage models like CLIP. This paper addresses the transductive zero-shot and few-shot CLIP classification challenge, in which inference is performed jointly across a mini-batch of unlabeled query samples, rather than treating each instance independently. We initially construct informative vision-text probability features, leading to a classification problem on the unit simplex set. Inspired by Expectation-Maximization (EM), our optimization-based classification objective models the data probability distribution for each class using a Dirichlet law. The minimization problem is then tackled with a novel block Majorization-Minimization algorithm, which simultaneously estimates the distribution parameters and class assignments. Extensive numerical experiments on 11 datasets underscore the benefits and efficacy of our batch inference approach. On zero-shot tasks with test batches of 75 samples, our approach yields near 20% improvement in ImageNet accuracy over CLIP's zero-shot performance. Additionally, we outperform state-of-the-art methods in the few-shot setting. The code is available at: https://github.com/SegoleneMartin/transductive-CLIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a33dbb08-bcfd-4162-a3da-cb8ed9b493b9Cited by top-tier papers20
- Boosting Vision-Language Models with TransductionMaxime Zanella, Benoît Gérin, Ismail Ben AyedNeurIPS 2024 · 42 citations
- Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIPYayuan Li, Jintao Guo, Lei Qi, Wenbin Li et al.AAAI 2025 · 9 citations
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation ModelsZhanxuan Hu, Qiyu Xu, Yu Duan, Yonghang Tai et al.CVPR 2026 · 6 citations
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningTianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICCV 2025 · 3 citations
- Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time TransductionJiazhen Huang, Zhiming Liu, Changhu Wang, Wei Ju et al.ICML 2026 · 2 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
Related papers
- TIM++: Transductive Information Maximization for Few-Shot CLIPYingping Li, Yutong Zou, Yunshi Huang, Changzhe Jiao et al.AAAI 2026
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 1 citation
- Generate, Transduct, Adapt: Iterative Transduction with VLMsOindrila Saha, Logan Lawrence, Grant Van Horn, Subhransu MajiICCV 2025 · 2 citations
- LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIPYunshi Huang, Fereshteh Shakeri, Jose Dolz, Malik Boudiaf et al.CVPR 2024
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
