A General Framework For Detecting Anomalous Inputs to DNN Classifiers
Jayaram Raghuram, Varun Chandrasekaran, Somesh Jha, Suman Banerjee
Abstract
Detecting anomalous inputs, such as adversarial and out-of-distribution (OOD) inputs, is critical for classifiers (including deep neural networks or DNNs) deployed in real-world applications. While prior works have proposed various methods to detect such anomalous samples using information from the internal layer representations of a DNN, there is a lack of consensus on a principled approach for the different components of such a detection method. As a result, often heuristic and one-off methods are applied for different aspects of this problem. We propose an unsupervised anomaly detection framework based on the internal DNN layer representations in the form of a meta-algorithm with configurable components. We proceed to propose specific instantiations for each component of the meta-algorithm based on ideas grounded in statistical testing and anomaly detection. We evaluate the proposed methods on well-known image classification datasets with strong adversarial attacks and OOD inputs, including an adaptive attack that uses the internal layer representations of the DNN (often not considered in prior work). Comparisons with five recently-proposed competing detection methods demonstrates the effectiveness of our method in detecting adversarial and OOD inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e5f78b1-b892-45d6-a6c2-8d4e15cc6797Cited by top-tier papers11
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying ThemFlorian TramèrICML 2022 · 82 citations
- Perturbation Learning Based Anomaly DetectionJinyu Cai, Jicong FanNeurIPS 2022 · 50 citations
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov et al.ICLR 2026 · 28 citations
- Concept-based Explanations for Out-of-Distribution DetectorsJihye Choi, Jayaram Raghuram, Ryan Feng, Jiefeng Chen et al.ICML 2023 · 18 citations
- Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionMaxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie C. K. Cheung et al.AAAI 2024 · 18 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Detecting Out-of-Distribution Examples with Gram MatricesChandramouli Shama Sastry, Sageev OoreICML 2020 · 275 citations
Related papers
- Detection of Out-of-Distribution Samples Using Binary Neuron Activation PatternsBartlomiej Olber, Krystian Radlak, Adam Popowicz, Michal Szczepankiewicz et al.CVPR 2023
- A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural NetworksMatan Haroush, Tzviel Frostig, Ruth Heller, Daniel SoudryICLR 2022 · 40 citations
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry et al.NeurIPS 2022 · 24 citations
- Generating Distributional Adversarial Examples to Evade Statistical DetectorsYigitcan Kaya, Muhammad Bilal Zafar, Sergül Aydöre, Nathalie Rauschmayr et al.ICML 2022 · 7 citations
- Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution DetectionJinglun Li, Xinyu Zhou, Pinxue Guo, Yixuan Sun et al.ICCV 2023 · 5 citations
