Certifiably Adversarially Robust Detection of Out-of-Distribution Data
Julian Bitterwolf, Alexander Meinke, Matthias Hein
Abstract
Deep neural networks are known to be overconfident when applied to out-of-distribution (OOD) inputs which clearly do not belong to any class. This is a problem in safety-critical applications since a reliable assessment of the uncertainty of a classifier is a key property, allowing the system to trigger human intervention or to transfer into a safe state. In this paper, we aim for certifiable worst case guarantees for OOD detection by enforcing not only low confidence at the OOD point but also in an -ball around it. For this purpose, we use interval bound propagation (IBP) to upper bound the maximal confidence in the -ball and minimize this upper bound during training time. We show that non-trivial bounds on the confidence for OOD data generalizing beyond the OOD dataset seen at training time are possible. Moreover, in contrast to certified adversarial robustness which typically comes with significant loss in prediction performance, certified guarantees for worst case OOD detection are possible without much loss in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 154 citations
- Boosting Out-of-distribution Detection with Typical FeaturesYao Zhu, Yuefeng Chen, Chuanlong Xie, Xiaodan Li et al.NeurIPS 2022 · 74 citations
- Negative Label Guided OOD Detection with Pretrained Vision-Language ModelsXue Jiang, Feng Liu, Zhen Fang, Hong Chen et al.ICLR 2024 · 73 citations
- Out-of-Distribution Detection with Negative PromptsJun Nie, Yonggang Zhang, Zhen Fang, Tongliang Liu et al.ICLR 2024 · 48 citations
- Watermarking for Out-of-distribution DetectionQizhou Wang, Feng Liu, Yonggang Zhang, Jing Zhang et al.NeurIPS 2022 · 44 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Sparse and Imperceivable Adversarial AttacksFrancesco Croce, Matthias HeinICCV 2019 · 228 citations
- Towards neural networks that provably know when they don't knowAlexander Meinke, Matthias HeinICLR 2020 · 151 citations
Related papers
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel et al.ICCV 2019 · 196 citations
- Provably robust classification of adversarial examples with detectionFatemeh Sheikholeslami, Ali Lotfi, J. Zico KolterICLR 2021 · 27 citations
- Understanding Certified Training with Interval Bound PropagationYuhao Mao, Mark Niklas Müller, Marc Fischer, Martin T. VechevICLR 2024 · 26 citations
- Certified Robustness for Deep Equilibrium Models via Interval Bound PropagationColin Wei, J. Zico KolterICLR 2022 · 22 citations
- On the Convergence of Certified Robust Training with Interval Bound PropagationYihan Wang, Zhouxing Shi, Quanquan Gu, Cho-Jui HsiehICLR 2022 · 11 citations
