Explaining Classifiers Using Adversarial Perturbations on the Perceptual Ball
Andrew Elliott, Stephen Law, Chris Russell
Abstract
We present a simple regularization of adversarial perturbations based upon the perceptual loss. While the resulting perturbations remain imperceptible to the human eye, they differ from existing adversarial perturbations in that they are semi-sparse alterations that highlight objects and regions of interest while leaving the background unaltered. As a semantically meaningful adverse perturbations, it forms a bridge between counterfactual explanations and adversarial perturbations in the space of images. We evaluate our approach on several standard explainability benchmarks, namely, weak localization, insertiondeletion, and the pointing game demonstrating that perceptually regularized counterfactuals are an effective explanation for image-based classifiers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 147 citations
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun et al.NeurIPS 2023 · 24 citations
- You Only Query Once: An Efficient Label-Only Membership Inference AttackYutong Wu, Han Qiu, Shangwei Guo, Jiwei Li et al.ICLR 2024 · 20 citations
- Do Perceptually Aligned Gradients Imply Robustness?Roy Ganz, Bahjat Kawar, Michael EladICML 2023 · 18 citations
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo et al.ICCV 2023 · 15 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke et al.ICCV 2019 · 160 citations
Related papers
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- IAP: Invisible Adversarial Patch Attack Through Perceptibility-Aware Localization and Perturbation OptimizationSubrat Kishore Dutta, Xiao ZhangICCV 2025
- Towards Large Yet Imperceptible Adversarial Image Perturbations With Perceptual Color DistanceZhengyu Zhao, Zhuoran Liu, Martha A. LarsonCVPR 2020
- Diffusion Visual Counterfactual ExplanationsMaximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias HeinNeurIPS 2022 · 124 citations
