Lune

NeurIPS2025Top-tier venue

Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection

Ignacio Meza De La Jara, Cristian Rodriguez Opazo, Damien Teney, Damith Ranasinghe, Ehsan Abbasnejad

2025Year
8Citations

Abstract

Out-of-distribution (OOD) detection is essential for reliably deploying machine learning models in the wild. Yet, most methods treat large pre-trained models as monolithic encoders and rely solely on their final-layer representations for detection. We challenge this wisdom. We reveal the intermediate layers of pre-trained models, shaped by residual connections that subtly transform input projections, can encode surprisingly rich and diverse signals for detecting distributional shifts. Importantly, to exploit latent representation diversity across layers, we introduce an entropybased criterion to automatically identify layers offering the most complementary information in a training-free setting-without access to OOD data. We show that selectively incorporating these intermediate representations can increase the accuracy of OOD detection by up to 10% in far-OOD and over 7% in near-OOD benchmarks compared to state-of-the-art training-free methods across various model architectures and training objectives. Our findings reveal a new avenue for OOD detection research and uncover the impact of various training objectives and model architectures on confidence-based OOD detection methods.

Recent work has leveraged large vision-language models (VLMs) such as CLIP [57] in attempts to address the problem. The methods enable zero-shot OOD detection by aligning image and text embeddings. However, the approaches implicitly treat these deep models as shallow because the information extracted for detection simply focus on the last layer. Indeed, deep neural networks typically rely on final-layer embeddings as compact, semantically rich representations of the input. But, a sole reliance on the last layer overlooks the neural structure through which they are obtained. Therefore, we challenge this widespread wisdom and propose exploring intermediate-layer representations. Interestingly, early work in convolutional models showed distinct functionality across layers [2] however pretrained vision transformers trained with diverse objectives-supervised, contrastive, or masked modeling-have not been thoroughly examined in this context. Understanding how in-and out-of-distribution semantics are distributed across depth may offer means for improving detection robustness and generalization. Our study revisits the problem of zero-shot OOD detection to ask: Can intermediate representations be systematically leveraged to improve detection performance?

To investigate, we conduct a comprehensive analysis across seven vision backbones, including CLIP, DINOv2 [54], MAE [23], MoCo-v3 [8], SiGLiP-v2 [61], Perception Encoder [4], and supervised 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 16c752d3-2b43-4782-b44c-8a93978198fd

Builds on34

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines