LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and Generalization
Jingzheng Li, Hailong Sun
Abstract
Pre-trained Vision-Language Models (VLMs) on large-scale image-text pairs, e.g., CLIP, have shown promising performance on zero-shot knowledge transfer. Recently, fine-tuning pre-trained VLMs to downstream few-shot classification with limited image annotation data yields significant gains. However, there are two limitations. First, most of the methods for fine-tuning VLMs only update newly added parameters while keeping the whole VLM frozen. Thus, it remains unclear how to directly update the VLM itself. Second, fine-tuning VLMs to a specific set of base classes would deteriorate the well-learned representation space such that the VLMs generalize poorly on novel classes. To address these issues, we first propose Layer-wise Fine-Tuning (LiFT) which achieves average gains of 3.9%, 4.3%, 4.2% and 4.5% on base classes under 2-, 4-, 8- and 16-shot respectively compared to the baseline CoOp over 11 datasets. Alternatively, we provide a parameter-efficient LiFT-Adapter exhibiting favorable performance while updating only 1.66% of total parameters. Further, we design scalable LiFT-NCD to identify both base classes and novel classes, which boosts the accuracy by an average of 5.01% over zero-shot generalization of CLIP, exploring the potential of VLMs in discovering novel classes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6df99a48-7b99-4895-af28-7add4b608aa6Cited by top-tier papers2
- Benchmarking In-the-Wild Multimodal Disease Recognition and A Versatile BaselineTianqi Wei, Zhi Chen, Zi Huang, Xin YuACM MM 2024 · 19 citations
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain AdaptationKerun Mi, Guoliang Kang, Guangyu Li, Lin Zhao et al.ACM MM 2025 · 1 citation
Related papers
- Learning to Learn Better Visual PromptsFengxiang Wang, Wanrong Huang, Shaowu Yang, Qi Fan et al.AAAI 2024 · 17 citations
- Learning Mask-aware CLIP Representations for Zero-Shot SegmentationSiyu Jiao, Yunchao Wei, Yaowei Wang, Yao Zhao et al.NeurIPS 2023 · 88 citations
- Adaptive Parameter Selection for Tuning Vision-Language ModelsYi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min HuCVPR 2025
- Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsYongjin Yang, Jongwoo Ko, Se-Young YunEMNLP 2024 · 1 citation
- CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentHaoyu Song, Li Dong, Weinan Zhang, Ting Liu et al.ACL 2022
