Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection
Ignacio Meza De La Jara, Cristian Rodriguez Opazo, Damien Teney, Damith Ranasinghe, Ehsan Abbasnejad
摘要
Out-of-distribution (OOD) detection is essential for reliably deploying machine learning models in the wild. Yet, most methods treat large pre-trained models as monolithic encoders and rely solely on their final-layer representations for detection. We challenge this wisdom. We reveal the intermediate layers of pre-trained models, shaped by residual connections that subtly transform input projections, can encode surprisingly rich and diverse signals for detecting distributional shifts. Importantly, to exploit latent representation diversity across layers, we introduce an entropybased criterion to automatically identify layers offering the most complementary information in a training-free setting-without access to OOD data. We show that selectively incorporating these intermediate representations can increase the accuracy of OOD detection by up to 10% in far-OOD and over 7% in near-OOD benchmarks compared to state-of-the-art training-free methods across various model architectures and training objectives. Our findings reveal a new avenue for OOD detection research and uncover the impact of various training objectives and model architectures on confidence-based OOD detection methods.
Recent work has leveraged large vision-language models (VLMs) such as CLIP [57] in attempts to address the problem. The methods enable zero-shot OOD detection by aligning image and text embeddings. However, the approaches implicitly treat these deep models as shallow because the information extracted for detection simply focus on the last layer. Indeed, deep neural networks typically rely on final-layer embeddings as compact, semantically rich representations of the input. But, a sole reliance on the last layer overlooks the neural structure through which they are obtained. Therefore, we challenge this widespread wisdom and propose exploring intermediate-layer representations. Interestingly, early work in convolutional models showed distinct functionality across layers [2] however pretrained vision transformers trained with diverse objectives-supervised, contrastive, or masked modeling-have not been thoroughly examined in this context. Understanding how in-and out-of-distribution semantics are distributed across depth may offer means for improving detection robustness and generalization. Our study revisits the problem of zero-shot OOD detection to ask: Can intermediate representations be systematically leveraged to improve detection performance?
To investigate, we conduct a comprehensive analysis across seven vision backbones, including CLIP, DINOv2 [54], MAE [23], MoCo-v3 [8], SiGLiP-v2 [61], Perception Encoder [4], and supervised 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Intermediate Layer Classifiers for OOD generalizationArnas Uselis, Seong Joon OhICLR 2025
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 被引用 9 次
- Exploring the Limits of Out-of-Distribution DetectionStanislav Fort, Jie Ren, Balaji LakshminarayananNeurIPS 2021 · 被引用 443 次
- X-Mahalanobis: Transformer Feature Mixing for Reliable OOD DetectionTong Wei, Bolin Wang, Jiang-Xin Shi, Yu-Feng Li 等NeurIPS 2025 · 被引用 9 次
- FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionXinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen 等ICCV 2025 · 被引用 1 次
