Your Out-of-Distribution Detection Method is Not Robust!
Mohammad Azizmalayeri, Arshia Soltani Moakhar, Arman Zarei, Reihaneh Zohrabi, Mohammad Taghi Manzuri, Mohammad Hossein Rohban
Abstract
Out-of-distribution (OOD) detection has recently gained substantial attention due to the importance of identifying out-of-domain samples in reliability and safety. Although OOD detection methods have advanced by a great deal, they are still susceptible to adversarial examples, which is a violation of their purpose. To mitigate this issue, several defenses have recently been proposed. Nevertheless, these efforts remained ineffective, as their evaluations are based on either small perturbation sizes, or weak attacks. In this work, we re-examine these defenses against an end-to-end PGD attack on in/out data with larger perturbation sizes, e.g. up to commonly used = 8/255 for the CIFAR-10 dataset. Surprisingly, almost all of these defenses perform worse than a random detection under the adversarial setting. Next, we aim to provide a robust OOD detection method. In an ideal defense, the training should expose the model to almost all possible adversarial perturbations, which can be achieved through adversarial training. That is, such training perturbations should based on both in-and out-of-distribution samples. Therefore, unlike OOD detection in the standard setting, access to OOD, as well as in-distribution, samples sounds necessary in the adversarial training setup. These tips lead us to adopt generative OOD detection methods, such as OpenGAN, as a baseline. We subsequently propose the Adversarially Trained Discriminator (ATD), which utilizes a pre-trained robust model to extract robust features, and a generator model to create OOD samples. We noted that, for the sake of training stability, in the adversarial training of the discriminator, one should attack real in-distribution as well as real outliers, but not generated outliers. Using ATD with CIFAR-10 and CIFAR-100 as the in-distribution data, we could significantly outperform all previous methods in the robust AUROC while maintaining high standard AUROC and classification accuracy. The code repository is available at https://github.com/rohban-lab/ATD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc26d95a-07d8-4c29-a127-52b82ccccb79Cited by top-tier papers6
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He et al.NeurIPS 2023 · 55 citations
- RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution SamplesHossein Mirzaei, Mohammad Jafari, Hamid Reza Dehbashi, Ali Ansari et al.ICML 2024 · 13 citations
- Robust One-Class Classification with Signed Distance Function using 1-Lipschitz Neural NetworksLouis Béthune, Paul Novello, Guillaume Coiffier, Thibaut Boissin et al.ICML 2023 · 12 citations
- Scanning Trojaned Models Using Out-of-Distribution SamplesHossein Mirzaei, Ali Ansari, Bahar Dibaei Nia, Mojtaba Nafez et al.NeurIPS 2024 · 6 citations
- Adversarially Robust Anomaly Detection through Spurious Negative Pair MitigationHossein Mirzaei, Mojtaba Nafez, Jafar Habibi, Mohammad Sabokrou et al.ICLR 2025
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
Related papers
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie et al.ICCV 2021 · 35 citations
- DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion ModelsRuiyuan Gao, Chenchen Zhao, Lanqing Hong, Qiang XuICCV 2023 · 29 citations
- Boosting the Adversarial Robustness of Graph Neural Networks: An OOD PerspectiveKuan Li, Yiwen Chen, Yang Liu, Jin Wang et al.ICLR 2024 · 13 citations
- On the Adversarial Robustness of Out-of-distribution Generalization ModelsXin Zou, Weiwei LiuNeurIPS 2023 · 10 citations
- OpenGAN: Open-Set Recognition via Open Data GenerationShu Kong, Deva RamananICCV 2021 · 7 citations
