Improved Visual Fine-tuning with Natural Language Supervision
Junyang Wang, Yuanhong Xu, Juhua Hu, Ming Yan, Jitao Sang, Qi Qian
Abstract
Fine-tuning a visual pre-trained model can leverage the semantic information from large-scale pre-training data and mitigate the over-fitting problem on downstream vision tasks with limited training examples. While the problem of catastrophic forgetting in pre-trained backbone has been extensively studied for fine-tuning, its potential bias from the corresponding pre-training task and data, attracts less attention. In this work, we investigate this problem by demonstrating that the obtained classifier after fine-tuning will be close to that induced by the pre-trained model. To reduce the bias in the classifier effectively, we introduce a reference distribution obtained from a fixed text classifier, which can help regularize the learned vision classifier. The proposed method, Text Supervised fine-tuning (TeS), is evaluated with diverse pre-trained vision models including ResNet and ViT, and text encoders including BERT and CLIP, on 11 downstream tasks. The consistent improvement with a clear margin over distinct scenarios confirms the effectiveness of our proposal. Code is available at https://github.com/idstcv/TeS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0eeccbbd-789d-4191-a3db-1fcc35b09eb5Cited by top-tier papers5
- Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIPQi Qian, Yuanhong Xu, Juhua HuNeurIPS 2023 · 34 citations
- Multi-Modal Proxy Learning Towards Personalized Visual Multiple ClusteringJiawei Yao, Qi Qian, Juhua HuCVPR 2024 · 19 citations
- Customized Multiple Clustering via Multi-Modal Subspace Proxy LearningJiawei Yao, Qi Qian, Juhua HuNeurIPS 2024 · 17 citations
- Enhancing Visual Continual Learning with Language-Guided SupervisionBolin Ni, Hongbo Zhao, Chenghao Zhang, Ke Hu et al.CVPR 2024 · 8 citations
- Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality CalibrationYiyang Chen, Tianyu Ding, Lei Wang, Jing Huo et al.CVPR 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan et al.AAAI 2026
- Debiased Fine-Tuning for Vision-Language Models by Prompt RegularizationBeier Zhu, Yulei Niu, Saeil Lee, Minhoe Hur et al.AAAI 2023 · 34 citations
- Global Knowledge Calibration for Fast Open-Vocabulary SegmentationKunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding et al.ICCV 2023 · 56 citations
- Anchor-based Robust Finetuning of Vision-Language ModelsJinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao et al.CVPR 2024
- Task Residual for Tuning Vision-Language ModelsTao Yu, Zhihe Lu, Xin Jin, Zhibo Chen et al.CVPR 2023
