AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
Yuwei Tang, Zhenyi Lin, Qilong Wang, Pengfei Zhu, Qinghua Hu
Abstract
Recently, pre-trained vision-language models (e.g., CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP, key factors on the effectiveness of existing methods have not been well studied, limiting further exploration of CLIP's potential in few-shot learning. In this paper, we first introduce a uni-fied formulation to analyze CLIP-based few-shot learning methods from a perspective of logit bias, which encourages us to learn an effective logit bias for further improving per-formance of CLIP-based few-shot learning methods. To this end, we disassemble three key components involved in computation of logit bias (i.e., logit features, logit predictor, and logit fusion) and empirically analyze the effect on per-formance of few-shot classification. Based on analysis of key components, this paper proposes a novel AMU-Tuning method to learn effective logit bias for CLIP-based few-shot classification. Specifically, our AMU-Tuning predicts logit bias by exploiting the appropriate Auxiliary features, which are fed into an efficient feature-initialized linear clas-sifier with Multi-branch training. Finally, an Uncertainty-based fusion is developed to incorporate logit bias into CLIP for few-shot classification. The experiments are con-ducted on several widely used benchmarks, and the re-sults show AMU-Tuning clearly outperforms its counter-parts while achieving state-of-the-art performance of CLIP-based few-shot learning without bells and whistles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIPYayuan Li, Jintao Guo, Lei Qi, Wenbin Li et al.AAAI 2025 · 9 citations
- Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li et al.CVPR 2026 · 7 citations
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li et al.CVPR 2026 · 6 citations
- Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local AlignmentYaze Zhao, Yixiong Zou, Yuhua Li, Ruixuan LiCVPR 2026 · 5 citations
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningTianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICCV 2025 · 3 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Logits DeConfusion with CLIP for Few-Shot LearningShuo Li, Fang Liu, Zehua Hao, Xinyi Wang et al.CVPR 2025
- CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentHaoyu Song, Li Dong, Weinan Zhang, Ting Liu et al.ACL 2022
- Knowledge-Aware Prompt Tuning for Generalizable Vision-Language ModelsBaoshuo Kan, Teng Wang, Wenpeng Lu, Xiantong Zhen et al.ICCV 2023 · 53 citations
- Vision-Language Model Fine-Tuning via Simple Parameter-Efficient ModificationMing Li, Jike Zhong, Chenxin Li, Liuzhuozheng Li et al.EMNLP 2024 · 18 citations
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
