Fooling Explanations in Text Classifiers
Adam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal Frossard
Abstract
State-of-the-art text classification models are becoming increasingly reliant on deep neural networks (DNNs). Due to their black-box nature, faithful and robust explanation methods need to accompany classifiers for deployment in real-life scenarios. However, it has been shown in vision applications that explanation methods are susceptible to local, imperceptible perturbations that can significantly alter the explanations without changing the predicted classes. We show here that the existence of such perturbations extends to text classifiers as well. Specifically, we introduceTextExplanationFooler (TEF), a novel explanation attack algorithm that alters text input samples imperceptibly so that the outcome of widely-used explanation methods changes considerably while leaving classifier predictions unchanged. We evaluate the performance of the attribution robustness estimation performance in TEF on five sequence classification datasets, utilizing three DNN architectures and three transformer architectures for each dataset. TEF can significantly decrease the correlation between unchanged and perturbed input attributions, which shows that all models and explanation methods are susceptible to TEF perturbations. Moreover, we evaluate how the perturbations transfer to other model architectures and attribution methods, and show that TEF perturbations are also effective in scenarios where the target model and explanation method are unknown. Finally, we introduce a semi-universal attack that is able to compute fast, computationally light perturbations with no knowledge of the attacked classifier nor explanation method. Overall, our work shows that explanations in text classifiers are very fragile and users need to carefully address their robustness before relying on them in critical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ff85f30-edd9-4929-a995-9571a689f984Cited by top-tier papers6
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 18 citations
- Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model PredictionsJingtan Wang, Xiaoqiang Lin, Rui Qiao, Chuan-Sheng Foo et al.ICML 2024 · 12 citations
- Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution ProtectionFan Wang, Adams Wai-Kin KongNeurIPS 2022 · 12 citations
- "Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial AttackChristopher Burger, Lingwei Chen, Thai LeEMNLP 2023 · 11 citations
Builds on1
Related papers
- Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target AttacksBohan Wang, Zewen Liu, Lu Lin, Hui Liu et al.ICML 2026 · 1 citation
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- ColorFool: Semantic Adversarial ColorizationAli Shahin Shamsabadi, Ricardo Sánchez-Matilla, Andrea CavallaroCVPR 2020
- On the Transferability of Adversarial Attacks against Neural Text ClassifierLiping Yuan, Xiaoqing Zheng, Yi Zhou, Cho-Jui Hsieh et al.EMNLP 2021 · 19 citations
- Corrupting Neuron Explanations of Deep Visual FeaturesDivyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei WengICCV 2023 · 3 citations
