Certifiably Adversarially Robust Detection of Out-of-Distribution Data
Julian Bitterwolf, Alexander Meinke, Matthias Hein
摘要
Deep neural networks are known to be overconfident when applied to out-of-distribution (OOD) inputs which clearly do not belong to any class. This is a problem in safety-critical applications since a reliable assessment of the uncertainty of a classifier is a key property, allowing the system to trigger human intervention or to transfer into a safe state. In this paper, we aim for certifiable worst case guarantees for OOD detection by enforcing not only low confidence at the OOD point but also in an -ball around it. For this purpose, we use interval bound propagation (IBP) to upper bound the maximal confidence in the -ball and minimize this upper bound during training time. We show that non-trivial bounds on the confidence for OOD data generalizing beyond the OOD dataset seen at training time are possible. Moreover, in contrast to certified adversarial robustness which typically comes with significant loss in prediction performance, certified guarantees for worst case OOD detection are possible without much loss in accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 被引用 154 次
- Boosting Out-of-distribution Detection with Typical FeaturesYao Zhu, Yuefeng Chen, Chuanlong Xie, Xiaodan Li 等NeurIPS 2022 · 被引用 74 次
- Negative Label Guided OOD Detection with Pretrained Vision-Language ModelsXue Jiang, Feng Liu, Zhen Fang, Hong Chen 等ICLR 2024 · 被引用 73 次
- Out-of-Distribution Detection with Negative PromptsJun Nie, Yonggang Zhang, Zhen Fang, Tongliang Liu 等ICLR 2024 · 被引用 48 次
- Watermarking for Out-of-distribution DetectionQizhou Wang, Feng Liu, Yonggang Zhang, Jing Zhang 等NeurIPS 2022 · 被引用 44 次
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Sparse and Imperceivable Adversarial AttacksFrancesco Croce, Matthias HeinICCV 2019 · 被引用 228 次
- Towards neural networks that provably know when they don't knowAlexander Meinke, Matthias HeinICLR 2020 · 被引用 151 次
相关 Paper
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel 等ICCV 2019 · 被引用 196 次
- Provably robust classification of adversarial examples with detectionFatemeh Sheikholeslami, Ali Lotfi, J. Zico KolterICLR 2021 · 被引用 27 次
- Understanding Certified Training with Interval Bound PropagationYuhao Mao, Mark Niklas Müller, Marc Fischer, Martin T. VechevICLR 2024 · 被引用 26 次
- Certified Robustness for Deep Equilibrium Models via Interval Bound PropagationColin Wei, J. Zico KolterICLR 2022 · 被引用 22 次
- On the Convergence of Certified Robust Training with Interval Bound PropagationYihan Wang, Zhouxing Shi, Quanquan Gu, Cho-Jui HsiehICLR 2022 · 被引用 11 次
