Adaptive Parameter Selection for Tuning Vision-Language Models
Yi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min Hu
Abstract
Vision-language models (VLMs) like CLIP have been widely used in various specific tasks. Parameter-efficient fine-tuning (PEFT) methods, such as prompt and adapter tuning, have become key techniques for adapting these models to specific domains. However, existing approaches rely on prior knowledge to manually identify the locations requiring fine-tuning. Adaptively selecting which parameters in VLMs should be tuned remains unexplored. In this paper, we propose CLIP with Adaptive Selective Tuning (CLIP-AST), which can be used to automatically select critical parameters in VLMs for fine-tuning for specific tasks. It opportunely leverages the adaptive learning rate in the optimizer and improves model performance without extra parameter overhead. We conduct extensive experiments on 13 benchmarks, such as ImageNet, Food101, Flowers102, etc, with different settings, including few-shot learning, base-to-novel class generalization, and out-ofdistribution. The results show that CLIP-AST consistently outperforms the original CLIP model as well as its variants and achieves state-of-the-art (SOTA) performance in all cases. For example, with the 16-shot learning, CLIP-AST surpasses GraphAdapter and PromptSRC by 3.56% and 2.20% in average accuracy on 11 datasets, respectively. Code will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ca47c3f-baac-4090-9dcd-70ef9a4458cdCited by top-tier papers1
Ask how each one uses itBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsZhengqin Xu, Zelin Peng, Xiaokang Yang, Wei ShenAAAI 2025 · 3 citations
- VioLET: Vision-Language Efficient Tuning with Collaborative Multi-modal GradientsYaoming Wang, Yuchen Liu, Xiaopeng Zhang, Jin Li et al.ACM MM 2023 · 2 citations
- Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPChen Huang, Skyler Seto, Samira Abnar, David Grangier et al.NeurIPS 2024 · 8 citations
- LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and GeneralizationJingzheng Li, Hailong SunACM MM 2023 · 5 citations
- Vision-Language Model Fine-Tuning via Simple Parameter-Efficient ModificationMing Li, Jike Zhong, Chenxin Li, Liuzhuozheng Li et al.EMNLP 2024 · 18 citations
