CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD Detection
Lin Zhu, Yifeng Yang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye
Abstract
Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair their ability to handle distribution shifts. In real-world scenarios, machine learning systems inevitably encounter both covariate shifts (e.g., changes in image styles) and semantic shifts (e.g., test-time unseen classes). This highlights the importance of enhancing out-of-distribution (OOD) generalization on covariate shifts and simultaneously detecting semantic-shifted unseen classes. Thus a critical but underexplored question arises: How to improve VL-PTMs' generalization ability to closed-set OOD data, while effectively detecting open-set unseen classes during fine-tuning? In this paper, we propose a novel objective function of OOD detection that also serves to improve OOD generalization. We show that minimizing the gradient magnitude of energy scores on training data leads to domain-consistent Hessians of classification loss, a strong indicator for OOD generalization revealed by theoretical analysis. Based on this finding, we have developed a unified fine-tuning framework that allows for concurrent optimization of both tasks. Extensive experiments have demonstrated the superiority of our method. The code is available at https://github.com/LinLLLL/CRoFT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce78ff4d-8c36-4086-91ec-50faa7b1d3baCited by top-tier papers4
- Exploring Channel-Aware Typical Features for Out-of-Distribution DetectionRundong He, Yue Yuan, Zhongyi Han, Fan Wang et al.AAAI 2024 · 8 citations
- Δ Energy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD GeneralizationLin Zhu, Yifeng Yang, Xinbing Wang, Qinying Gu et al.NeurIPS 2025 · 2 citations
- Detecting Out-of-Distribution Through the Lens of Neural CollapseLitian Liu, Yao QinCVPR 2025
- Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution DetectionXiang Fang, Arvind Easwaran, Blaise GenestICML 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
Related papers
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon et al.ICML 2023 · 67 citations
- UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language ModelingYuchuan Li, Azadeh Motamedi, Hyock Ju Kwon, Chul B Park et al.CVPR 2026
- Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningKun Ding, Haojian Zhang, Qiang Yu, Ying Wang et al.AAAI 2024 · 8 citations
- Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language ModelsYuanwei Hu, Bo Peng, Yadan Luo, zhen fang et al.ICML 2026
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 9 citations
