Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
Yuanwei Hu, Bo Peng, Yadan Luo, zhen fang, Ling Chen, Jie Lu
Abstract
Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected inputs from unknown classes. Recent progress in pre-trained vision-language models (VLMs) has enabled zero-shot OOD detection without access to in-distribution (ID) training data; in this setting, existing methods commonly treat text embeddings of class names as class prototypes. In this paper, we challenge this widely adopted "text-as-prototype" paradigm by theoretically showing that off-the-shelf textual prototypes are generally misaligned with the optimal visual prototypes, yielding an intrinsic modality gap that cannot be eliminated by prompt engineering alone. To mitigate this gap under the post-hoc constraint, this paper presents an online pseudo-supervised framework that directly learns class prototypes in the visual feature space using unlabeled test-time data streams and soft predictions from the pretrained VLMs. We provide theoretical guarantees for the convergence of the online optimization procedure. Extensive experiments empirically manifest that our method achieves a new state of the art across a variety of OOD detection setups. Code is available at here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69e6a136-562a-4c8b-8f7c-fb5cf0153752Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language ModelsBo Peng, Jie Lu, Guangquan Zhang, Zhen FangKDD 2026 · 1 citation
- Delving into Out-of-Distribution Detection with Vision-Language RepresentationsYifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun et al.NeurIPS 2022 · 308 citations
- Negative Label Guided OOD Detection with Pretrained Vision-Language ModelsXue Jiang, Feng Liu, Zhen Fang, Hong Chen et al.ICLR 2024 · 73 citations
- Dynamic Multimodal Prototype Learning in Vision-Language ModelsXingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li et al.ICCV 2025
- TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language ModelsJinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu et al.CVPR 2026 · 2 citations
