Anchor-based Robust Finetuning of Vision-Language Models
Jinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao, Ke Yan, Shouhong Ding, Yuan Gao, Gui-Song Xia
Abstract
We aim at finetuning a vision-language model without hurting its out-of-distribution (OOD) generalization. We address two types of OOD generalization, i.e., i) domain shift such as natural to sketch images, and ii) zero-shot capability to recognize the category that was not contained in the finetune data. Arguably, the diminished OOD generalization after finetuning stems from the excessively simplified finetuning target, which only provides the class information, such as "a photo of a [CLASS]". This is distinct from the process in that CLIP was pretrained, where there is abundant text supervision with rich semantic information. Therefore, we propose to compensate for the finetune process using auxiliary supervision with rich semantic information, which acts as anchors to preserve the OOD generalization. Specifically, two types of anchors are elaborated in our method, including i) text-compensated anchor which uses the images from the finetune set but enriches the text supervision from a pretrained captioner, ii) image-text-pair anchor which is retrieved from the dataset similar to pretraining data of CLIP according to the downstream task, associating with the original CLIP text with rich semantics. Those anchors are utilized as auxiliary semantic information to maintain the original feature space of CLIP, thereby preserving the OOD generalization capabilities. Comprehensive experiments demonstrate that our method achieves in-distribution performance akin to conventional finetuning while attaining new state-of-the-art results on domain shift and zero-shot learning benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsChangdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han et al.NeurIPS 2024 · 49 citations
- SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained ModelsLinglan Zhao, Xuerui Zhang, Ke Yan, Shouhong Ding et al.NeurIPS 2024 · 22 citations
- Robust Fine-tuning of Zero-shot Models via Variance ReductionBeier Zhu, Jiequan Cui, Hanwang ZhangNeurIPS 2024 · 9 citations
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICLR 2026 · 1 citation
- TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label NoiseTengfei Ma, Weiran Pan, Wei WeiCVPR 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPChen Huang, Skyler Seto, Samira Abnar, David Grangier et al.NeurIPS 2024 · 8 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentHaoyu Song, Li Dong, Weinan Zhang, Ting Liu et al.ACL 2022
- Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane et al.AAAI 2026
