SoK: Explainable Machine Learning in Adversarial Environments
Maximilian Noppel, Christian Wressnegger
Abstract
Modern deep learning methods have long been considered black boxes due to the lack of insights into their decision-making process. However, recent advances in explainable machine learning have turned the tables. Post-hoc explanation methods enable precise relevance attribution of input features for otherwise opaque models such as deep neural networks. This progression has raised expectations that these techniques can uncover attacks against learning-based systems such as adversarial examples or neural backdoors. Unfortunately, current methods are not robust against manipulations themselves. In this paper, we set out to systematize attacks against post-hoc explanation methods to lay the groundwork for developing more robust explainable machine learning. If explanation methods cannot be misled by an adversary, they can serve as an effective tool against attacks, marking a turning point in adversarial machine learning. We present a hierarchy of explanation-aware robustness notions and relate existing defenses to it. In doing so, we uncover synergies, research gaps, and future directions toward more reliable explanations robust against manipulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ed67b4f-ada6-4900-a231-be5bd5dc3953Cited by top-tier papers2
- SoK: Unintended Interactions among Machine Learning Defenses and RisksVasisht Duddu, Sebastian Szyller, N. AsokanS&P 2024 · 6 citations
- EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace MethodYanting Wang, Jinyuan JiaICLR 2026
Builds on52
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
Related papers
- Disguising Attacks with Explanation-Aware BackdoorsMaximilian Noppel, Lukas Peter, Christian WressneggerS&P 2023
- XRand: Differentially Private Defense against Explanation-Guided AttacksTruc D. T. Nguyen, Phung Lai, Hai Phan, My T. ThaiAAAI 2023 · 22 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong et al.ICLR 2024 · 3 citations
