Lune

NeurIPS2025顶会

Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection

Ignacio Meza De La Jara, Cristian Rodriguez Opazo, Damien Teney, Damith Ranasinghe, Ehsan Abbasnejad

2025年份
8被引次数

摘要

Out-of-distribution (OOD) detection is essential for reliably deploying machine learning models in the wild. Yet, most methods treat large pre-trained models as monolithic encoders and rely solely on their final-layer representations for detection. We challenge this wisdom. We reveal the intermediate layers of pre-trained models, shaped by residual connections that subtly transform input projections, can encode surprisingly rich and diverse signals for detecting distributional shifts. Importantly, to exploit latent representation diversity across layers, we introduce an entropybased criterion to automatically identify layers offering the most complementary information in a training-free setting-without access to OOD data. We show that selectively incorporating these intermediate representations can increase the accuracy of OOD detection by up to 10% in far-OOD and over 7% in near-OOD benchmarks compared to state-of-the-art training-free methods across various model architectures and training objectives. Our findings reveal a new avenue for OOD detection research and uncover the impact of various training objectives and model architectures on confidence-based OOD detection methods.

Recent work has leveraged large vision-language models (VLMs) such as CLIP [57] in attempts to address the problem. The methods enable zero-shot OOD detection by aligning image and text embeddings. However, the approaches implicitly treat these deep models as shallow because the information extracted for detection simply focus on the last layer. Indeed, deep neural networks typically rely on final-layer embeddings as compact, semantically rich representations of the input. But, a sole reliance on the last layer overlooks the neural structure through which they are obtained. Therefore, we challenge this widespread wisdom and propose exploring intermediate-layer representations. Interestingly, early work in convolutional models showed distinct functionality across layers [2] however pretrained vision transformers trained with diverse objectives-supervised, contrastive, or masked modeling-have not been thoroughly examined in this context. Understanding how in-and out-of-distribution semantics are distributed across depth may offer means for improving detection robustness and generalization. Our study revisits the problem of zero-shot OOD detection to ask: Can intermediate representations be systematically leveraged to improve detection performance?

To investigate, we conduct a comprehensive analysis across seven vision backbones, including CLIP, DINOv2 [54], MAE [23], MoCo-v3 [8], SiGLiP-v2 [61], Perception Encoder [4], and supervised 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper34

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖