Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
Jiang-Xin Shi, Chi Zhang, Tong Wei, Yufeng Li
Abstract
Pre-trained vision-language models like CLIP have shown powerful zero-shot inference ability via image-text matching and prove to be strong few-shot learners in various downstream tasks. However, in real-world scenarios, adapting CLIP to downstream tasks may encounter the following challenges: 1) data may exhibit long-tailed data distributions and might not have abundant samples for all the classes; 2) There might be emerging tasks with new classes that contain no samples at all. To overcome them, we propose a novel framework to achieve efficient and long-tailed generalization, which can be termed as Candle. During the training process, we propose compensating logit-adjusted loss to encourage large margins of prototypes and alleviate imbalance both within the base classes and between the base and new classes. For efficient adaptation, we treat the CLIP model as a black box and leverage the extracted features to obtain visual and textual prototypes for prediction. To make full use of multi-modal information, we also propose cross-modal attention to enrich the features from both modalities. For effective generalization, we introduce virtual prototypes for new classes to make up for their lack of training images. Candle achieves state-of-the-art performance over extensive experiments on 11 diverse datasets while substantially reducing the training time, demonstrating the superiority of our approach. The source code is available at https://github.com/shijxcs/Candle . CCS CONCEPTS • Computing methodologies → Supervised learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9d75578-1988-4ff0-8905-413a410b5b8aCited by top-tier papers6
- DeCoOp: Robust Prompt Tuning with Out-of-Distribution DetectionZhi Zhou, Ming Yang, Jiang-Xin Shi, Lan-Zhe Guo et al.ICML 2024 · 14 citations
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan et al.AAAI 2026
- Long-tailed Test-Time Adaptation for Vision-Language ModelsXucong Wang, Zhe Zhao, Zekun Wang, Xiaofeng Cao et al.ICLR 2026
- Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language ModelsBiao Chen, Lin Zuo, Mengmeng Jing, Kunbin He et al.AAAI 2026
- Adaptive Token Refinement in Long-Tailed Large Vision-Language Models Fine-TuningWenjun Miao, Mingda Li, Yanchao Hao, Zheng WeiICML 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentXi Yu, Shinjae Yoo, Yuewei LinNeurIPS 2024 · 36 citations
- Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsYongjin Yang, Jongwoo Ko, Se-Young YunEMNLP 2024 · 1 citation
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- Probabilistic Prompt Distribution Learning for Animal Pose EstimationJiyong Rao, Brian Nlong Zhao, Yu WangCVPR 2025
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
