Your Out-of-Distribution Detection Method is Not Robust!
Mohammad Azizmalayeri, Arshia Soltani Moakhar, Arman Zarei, Reihaneh Zohrabi, Mohammad Taghi Manzuri, Mohammad Hossein Rohban
摘要
Out-of-distribution (OOD) detection has recently gained substantial attention due to the importance of identifying out-of-domain samples in reliability and safety. Although OOD detection methods have advanced by a great deal, they are still susceptible to adversarial examples, which is a violation of their purpose. To mitigate this issue, several defenses have recently been proposed. Nevertheless, these efforts remained ineffective, as their evaluations are based on either small perturbation sizes, or weak attacks. In this work, we re-examine these defenses against an end-to-end PGD attack on in/out data with larger perturbation sizes, e.g. up to commonly used = 8/255 for the CIFAR-10 dataset. Surprisingly, almost all of these defenses perform worse than a random detection under the adversarial setting. Next, we aim to provide a robust OOD detection method. In an ideal defense, the training should expose the model to almost all possible adversarial perturbations, which can be achieved through adversarial training. That is, such training perturbations should based on both in-and out-of-distribution samples. Therefore, unlike OOD detection in the standard setting, access to OOD, as well as in-distribution, samples sounds necessary in the adversarial training setup. These tips lead us to adopt generative OOD detection methods, such as OpenGAN, as a baseline. We subsequently propose the Adversarially Trained Discriminator (ATD), which utilizes a pre-trained robust model to extract robust features, and a generator model to create OOD samples. We noted that, for the sake of training stability, in the adversarial training of the discriminator, one should attack real in-distribution as well as real outliers, but not generated outliers. Using ATD with CIFAR-10 and CIFAR-100 as the in-distribution data, we could significantly outperform all previous methods in the robust AUROC while maintaining high standard AUROC and classification accuracy. The code repository is available at https://github.com/rohban-lab/ATD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He 等NeurIPS 2023 · 被引用 55 次
- RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution SamplesHossein Mirzaei, Mohammad Jafari, Hamid Reza Dehbashi, Ali Ansari 等ICML 2024 · 被引用 13 次
- Robust One-Class Classification with Signed Distance Function using 1-Lipschitz Neural NetworksLouis Béthune, Paul Novello, Guillaume Coiffier, Thibaut Boissin 等ICML 2023 · 被引用 12 次
- Scanning Trojaned Models Using Out-of-Distribution SamplesHossein Mirzaei, Ali Ansari, Bahar Dibaei Nia, Mojtaba Nafez 等NeurIPS 2024 · 被引用 6 次
- Adversarially Robust Anomaly Detection through Spurious Negative Pair MitigationHossein Mirzaei, Mojtaba Nafez, Jafar Habibi, Mohammad Sabokrou 等ICLR 2025
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
相关 Paper
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie 等ICCV 2021 · 被引用 35 次
- DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion ModelsRuiyuan Gao, Chenchen Zhao, Lanqing Hong, Qiang XuICCV 2023 · 被引用 29 次
- Boosting the Adversarial Robustness of Graph Neural Networks: An OOD PerspectiveKuan Li, Yiwen Chen, Yang Liu, Jin Wang 等ICLR 2024 · 被引用 13 次
- On the Adversarial Robustness of Out-of-distribution Generalization ModelsXin Zou, Weiwei LiuNeurIPS 2023 · 被引用 10 次
- OpenGAN: Open-Set Recognition via Open Data GenerationShu Kong, Deva RamananICCV 2021 · 被引用 7 次
