Delving into Out-of-Distribution Detection with Vision-Language Representations
Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun, Wei Li, Yixuan Li
Abstract
Recognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or language), leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of OOD detection from a single-modal to a multi-modal regime. Particularly, we propose Maximum Concept Matching (MCM), a simple yet effective zero-shot OOD detection method based on aligning visual features with textual concepts. We contribute in-depth analysis and theoretical insights to understand the effectiveness of MCM. Extensive experiments demonstrate that MCM achieves superior performance on a wide variety of real-world tasks. MCM with vision-language features outperforms a common baseline with pure visual features on a hard OOD task with semantically similar classes by 13.1% (AUROC). Code is available at https://github.com/ deeplearning-wisc/MCM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3187af7b-2f7b-47fe-bccf-eda63f894b64Cited by top-tier papers111
- LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt LearningAtsuyuki Miyai, Qing Yu, Go Irie, Kiyoharu AizawaNeurIPS 2023 · 174 citations
- CLIPN for Zero-Shot OOD Detection: Teaching CLIP to Say NoHualiang Wang, Yi Li, Huifeng Yao, Xiaomeng LiICCV 2023 · 171 citations
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 154 citations
- A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)Weijie Tu, Weijian Deng, Tom GedeonNeurIPS 2023 · 74 citations
- Negative Label Guided OOD Detection with Pretrained Vision-Language ModelsXue Jiang, Feng Liu, Zhen Fang, Hong Chen et al.ICLR 2024 · 73 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language ModelsYuanwei Hu, Bo Peng, Yadan Luo, zhen fang et al.ICML 2026
- Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIPSepideh Esmaeilpour, Bing Liu, Eric Robertson, Lei ShuAAAI 2022 · 219 citations
- DualCnst: Enhancing Zero-Shot Out-of-Distribution Detection via Text-Image Consistency in Vision-Language ModelsFayi Le, Wenwu He, Chentao Cao, Dong Liang et al.NeurIPS 2025 · 2 citations
- UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language ModelingYuchuan Li, Azadeh Motamedi, Hyock Ju Kwon, Chul B Park et al.CVPR 2026
- Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution DetectionKaixiang Chen, Pengfei Fang, Hui XueAAAI 2026
