Provably Safeguarding a Classifier from OOD and Adversarial Samples
Nicolas Atienza, Johanne Cohen, Christophe Labreuche, Michèle Sebag
Abstract
This paper aims to transform a trained classifier into an abstaining classifier, suchthat the latter is provably protected from out-of-distribution and adversarial samples. The proposed Sample-efficient Probabilistic Detection using Extreme ValueTheory (SPADE) approach relies on a Generalized Extreme Value (GEV) modelof the training distribution in the latent space of the classifier. Under mild assumptions, this GEV model allows for formally characterizing out-of-distributionand adversarial samples and rejecting them. Empirical validation of the approachis conducted on various neural architectures (ResNet, VGG, and Vision Transformer) and considers medium and large-sized datasets (CIFAR-10, CIFAR-100,and ImageNet). The results show the stability and frugality of the GEV model anddemonstrate SPADE’s efficiency compared to the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02a2ffc7-dbf8-455e-809d-f954dbf9070eBuilds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
Related papers
- DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion ModelsRuiyuan Gao, Chenchen Zhao, Lanqing Hong, Qiang XuICCV 2023 · 29 citations
- SPADE: A Spectral Method for Black-Box Adversarial Robustness EvaluationWuxinlin Cheng, Chenhui Deng, Zhiqiang Zhao, Yaohui Cai et al.ICML 2021 · 24 citations
- Unsupervised Out-of-Domain Detection via Pre-trained TransformersKeyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng et al.ACL 2021
- REST: Performance Improvement of a Black Box Model via RL-Based Spatial TransformationJae-Myung Kim, Hyungjin Kim, Chanwoo Park, Jungwoo LeeAAAI 2020
- GEN: Pushing the Limits of Softmax-Based Out-of-Distribution DetectionXixi Liu, Yaroslava Lochman, Christopher ZachCVPR 2023
