Provably Adversarially Robust Detection of Out-of-Distribution Data (Almost) for Free
Alexander Meinke, Julian Bitterwolf, Matthias Hein
Abstract
The application of machine learning in safety-critical systems requires a reliable assessment of uncertainty. However, deep neural networks are known to produce highly overconfident predictions on out-of-distribution (OOD) data. Even if trained to be non-confident on OOD data, one can still adversarially manipulate OOD data so that the classifier again assigns high confidence to the manipulated samples. We show that two previously published defenses can be broken by better adapted attacks, highlighting the importance of robustness guarantees around OOD data. Since the existing method for this task is hard to train and significantly limits accuracy, we construct a classifier that can simultaneously achieve provably adversarially robust OOD detection and high clean accuracy. Moreover, by slightly modifying the classifier's architecture our method provably avoids the asymptotic overconfidence problem of standard neural networks. We provide code for all our experiments. †
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1546c5dc-4cee-4d3f-800b-b7331f5d2aa8Cited by top-tier papers6
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 154 citations
- RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution SamplesHossein Mirzaei, Mohammad Jafari, Hamid Reza Dehbashi, Ali Ansari et al.ICML 2024 · 13 citations
- Open Set Label Shift with Test Time Out-of-Distribution ReferenceChangkun Ye, Russell Tsuchida, Lars Petersson, Nick BarnesCVPR 2025
- Adversarially Robust Out-of-Distribution Detection Using Lyapunov-Stabilized EmbeddingsHossein Mirzaei, Mackenzie W. MathisICLR 2025
- Inside-Out: Measuring Generalization in Vision Transformers Through Inner WorkingsYunxiang Peng, Mengmeng Ma, Ziyu Yao, Xi PengCVPR 2026
Builds on11
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
Related papers
- Towards neural networks that provably know when they don't knowAlexander Meinke, Matthias HeinICLR 2020 · 151 citations
- Certifiably Adversarially Robust Detection of Out-of-Distribution DataJulian Bitterwolf, Alexander Meinke, Matthias HeinNeurIPS 2020 · 91 citations
- iDECODe: In-Distribution Equivariance for Conformal Out-of-Distribution DetectionRamneet Kaur, Susmit Jha, Anirban Roy, Sangdon Park et al.AAAI 2022 · 53 citations
- Self-Supervised Learning for Generalizable Out-of-Distribution DetectionSina Mohseni, Mandar Pitale, J. B. S. Yadawa, Zhangyang WangAAAI 2020 · 229 citations
- Enhancing Adversarial Robustness with Conformal Prediction: A Framework for Guaranteed Model ReliabilityJie Bao, Chuangyin Dang, Rui Luo, Hanwei Zhang et al.ICML 2025
