Lune

NeurIPS2025Top-tier venue

Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection

Reihaneh Zohrabi, Hosein Hasani, Mahdieh Soleymani Baghshah, Anna Rohrbach, Marcus Rohrbach, Mohammad Hossein Rohban

2025Year
6Citations
1Top-tier citations

Abstract

Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications, where they frequently face data distributions unseen during training. Despite progress, existing methods are often vulnerable to spurious correlations that mislead models and compromise robustness. To address this, we propose SPROD, a novel prototype-based OOD detection approach that explicitly addresses the challenge posed by unknown spurious correlations. Our post-hoc method refines class prototypes to mitigate bias from spurious features without additional data or hyperparameter tuning, and is broadly applicable across diverse backbones and OOD detection settings. We conduct a comprehensive spurious correlation OOD detection benchmarking, comparing our method against existing approaches and demonstrating its superior performance across challenging OOD datasets, such as CelebA, Waterbirds, UrbanCars, Spurious Imagenet, and the newly introduced Animals MetaCoCo. On average, SPROD improves AUROC by 4.8% and FPR@95 by 9.4% over the second best.

• This work conducts and introduces comprehensive benchmarking across multiple SP-OOD datasets, including the newly introduced Animals MetaCoCo, a realistic, multiclass dataset with diverse spurious attributes.

• Finally, our study sheds new light on key factors influencing SP-OOD detection, such as the impact of backbone fine-tuning and the choice of scoring mechanisms.

2 Related Work OOD detection methods can be categorized into training-time and post-hoc approaches [37]. Trainingtime methods leverage auxiliary OOD samples (Outlier Exposure) [12, 38, 39] or apply regularization [40-43] to enhance OOD detection. Post-hoc methods, in contrast, derive OOD scores from base classifiers without modifying training [37]. Overall, post-hoc methods offer simplicity and competitive performance [37], making them practical under limited data or training resources. Among post-hoc methods, several approaches apply transformations to model logits to derive OOD scores. MSP [7] uses the maximum softmax probability, the energy-based method [13] computes the log-sum-exp of logits, MLS [44] uses the maximum logit and introduces KL Matching (KLM) based on KL divergence, and GEN [45] employs generalized entropy of softmax outputs. Another class of post-hoc methods detects OOD samples via feature-space distances. MDS [11] fits class-conditional Gaussians to pre-logit features and computes Mahalanobis distances, refined by RMDS [46] with an unconditional Gaussian on ID data. KNN [47] uses distances to nearest ID samples. SHE [48] scores samples by their distance to stored ID feature templates. NNGuide [49] leverages nearest-neighbor guidance to adjust test features toward the ID manifold. Relation [50] constructs a graph over training embeddings and detects outliers via relational anomalies. NECO [51] scores samples by their feature alignment with class weight vectors, leveraging neural collapse geometry. SCALE [52] separates ID and OOD samples by scaling penultimate-layer activations. FDBD [53] measures features' regularized mean distance to the classifier's decision boundaries. NCI [54] scores samples by their distance to class weight vectors, filtered by feature norms. Prototype-based methods shape class representations for distance-based OOD scoring. Classical approaches like MDS [11] and its variants [46] are closely related, as they model each class by a centroid in feature space, optionally using class-conditional covariances to compute Mahalanobis distances. Recent works extend this via explicit training objectives. CIDER [55] learns hyperspherical embeddings by jointly enforcing intra-class compactness and inter-class dispersion, thereby improving ID and OOD separability. PALM [56] represents each class as a mixture of learnable prototypes and optimizes a maximum-likelihood and contrastive objective, updating prototypes and backbone features jointly during training. PROWL [57] also leverages prototype representations, but for pixel-level OOD detection in segmentation. While these methods share a prototypical framework, SPROD differs as a post-hoc method operating on pretrained backbones and is explicitly designed to mitigate the negative effects of unknown spurious correlations. Appendix H further analyzes a variant, SPROD-KMeans, which connects to mixture-of-prototypes ideas while remaining fully post-hoc. A few methods combine information from both feature and logit spaces. ReAct [58] thresholds activations before applying energy-based scoring. ViM [59] adds a virtual logit from the residual norm between input features and the ID subspace and applies softmax over extended logits. ASH [60] prunes high-magnitude activations and rescales remaining features before logit computation, improving energy-based OOD separability. Some methods also exploit gradient space for OOD scoring [61, 62]. GradNorm [61] computes the KL divergence to a uniform

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on39

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines