Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
Ming Li, Jike Zhong, Chenxin Li, Liuzhuozheng Li, Nie Lin, Masashi Sugiyama
摘要
Recent advances in fine-tuning Vision-Language Models (VLMs) have witnessed the success of prompt tuning and adapter tuning, while the classic model fine-tuning on inherent parameters seems to be overlooked. It is believed that fine-tuning the parameters of VLMs with few-shot samples corrupts the pre-trained knowledge since fine-tuning the CLIP model even degrades performance. In this paper, we revisit this viewpoint, and propose a new perspective: fine-tuning the specific parameters instead of all will uncover the power of classic model fine-tuning on VLMs. Through our meticulous study, we propose ClipFit, a simple yet effective method to fine-tune CLIP without introducing any overhead of extra parameters. We demonstrate that by only fine-tuning the specific bias terms and normalization layers, ClipFit can improve the performance of zero-shot CLIP by 7.27% average harmonic mean accuracy. Lastly, to understand how fine-tuning in CLIPFit affects the pre-trained models, we conducted extensive experimental analyses w.r.t. changes in internal parameters and representations. We found that low-level text bias layers and the first layer normalization layer change much more than other layers. The code is available at https://github.com/minglllli/CLIPFit .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Approximate Domain Unlearning for Vision-Language ModelsKodai Kawamura, Yuta Goto, Rintaro Yanagi, Hirokatsu Kataoka 等NeurIPS 2025 · 被引用 7 次
- To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-TuningMing Li, Jike Zhong, Shitian Zhao, Yuxiang Lai 等NeurIPS 2025 · 被引用 5 次
- NormFit: A Lightweight Solution for Few-Shot Federated Learning with Non-IID DataAzadeh Motamedi, Jae-Mo Kang, Il-Min KimNeurIPS 2025 · 被引用 1 次
- InfoBridge: Balanced Multimodal Integration through Conditional Dependency ModelingChenxin Li, Yifan Liu, Panwang Pan, Hengyu Liu 等ICCV 2025 · 被引用 1 次
- EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering BenchmarkMing Li, Jike Zhong, Tianle Chen, Yuxiang Lai 等CVPR 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Adaptive Parameter Selection for Tuning Vision-Language ModelsYi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min HuCVPR 2025
- LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and GeneralizationJingzheng Li, Hailong SunACM MM 2023 · 被引用 5 次
- Understanding Zero-shot Adversarial Robustness for Large-Scale ModelsChengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang 等ICLR 2023 · 被引用 10 次
- One Last Attention for Your Vision-Language ModelLiang Chen, Ghazi Shazan Ahmad, Tianjun Yao, Lingqiao Liu 等ICCV 2025 · 被引用 1 次
- Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language ModelsKecheng Zheng, Wei Wu, Ruili Feng, Kai Zhu 等ICCV 2023 · 被引用 13 次
