Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
Jian Liang, Lijun Sheng, Zhengbo Wang, Ran He, Tieniu Tan
Abstract
The emergence of vision-language models, such as CLIP, has spurred a significant research effort towards their application for downstream supervised learning tasks. Although some previous studies have explored the unsupervised fine-tuning of CLIP, they often rely on prior knowledge in the form of class names associated with ground truth labels. This paper explores a realistic unsupervised fine-tuning scenario, considering the presence of out-of-distribution samples from unknown classes within the unlabeled data. In particular, we focus on simultaneously enhancing out-of-distribution detection and the recognition of instances associated with known classes. To tackle this problem, we present a simple, efficient, and effective approach called Universal Entropy Optimization (UEO). UEO leverages sample-level confidence to approximately minimize the conditional entropy of confident instances and maximize the marginal entropy of less confident instances. Apart from optimizing the textual prompt, UEO incorporates optimization of channel-wise affine transformations within the visual branch of CLIP. Extensive experiments across 15 domains and 4 different types of prior knowledge validate the effectiveness of UEO compared to baseline methods. The code is publicly available at https://github.com/tim-learn/UEO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Cooperative Pseudo Labeling for Unsupervised Federated ClassificationKuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang et al.ICCV 2025 · 1 citation
- COME: Test-time Adaption by Conservatively Minimizing EntropyQingyang Zhang, Yatao Bian, Xinke Kong, Peilin Zhao et al.ICLR 2025
- Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language ModelsYongguang Li, Jindong Li, Qi Wang, Qianli Xing et al.AAAI 2026
- LoRA-Pro: Are Low-Rank Adapters Properly Optimized?Zhengbo Wang, Jian Liang, Ran He, Zilei Wang et al.ICLR 2025
- Generalizing Vision-Language Models with Dedicated Prompt GuidanceXinyao Li, Yinjie Min, Hongbo Chen, Zhekai Du et al.AAAI 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
Related papers
- Unknown Text Learning for Clip-Based Few-Shot Open-Set RecognitionRui Ma, Qilong Wang, Bing Cao, Qinghua Hu et al.ICCV 2025 · 1 citation
- DeCoOp: Robust Prompt Tuning with Out-of-Distribution DetectionZhi Zhou, Ming Yang, Jiang-Xin Shi, Lan-Zhe Guo et al.ICML 2024 · 14 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt LearningAtsuyuki Miyai, Qing Yu, Go Irie, Kiyoharu AizawaNeurIPS 2023 · 174 citations
- Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationJameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak et al.NeurIPS 2023 · 147 citations
