Fooling Explanations in Text Classifiers
Adam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal Frossard
摘要
State-of-the-art text classification models are becoming increasingly reliant on deep neural networks (DNNs). Due to their black-box nature, faithful and robust explanation methods need to accompany classifiers for deployment in real-life scenarios. However, it has been shown in vision applications that explanation methods are susceptible to local, imperceptible perturbations that can significantly alter the explanations without changing the predicted classes. We show here that the existence of such perturbations extends to text classifiers as well. Specifically, we introduceTextExplanationFooler (TEF), a novel explanation attack algorithm that alters text input samples imperceptibly so that the outcome of widely-used explanation methods changes considerably while leaving classifier predictions unchanged. We evaluate the performance of the attribution robustness estimation performance in TEF on five sequence classification datasets, utilizing three DNN architectures and three transformer architectures for each dataset. TEF can significantly decrease the correlation between unchanged and perturbed input attributions, which shows that all models and explanation methods are susceptible to TEF perturbations. Moreover, we evaluate how the perturbations transfer to other model architectures and attribution methods, and show that TEF perturbations are also effective in scenarios where the target model and explanation method are unknown. Finally, we introduce a semi-universal attack that is able to compute fast, computationally light perturbations with no knowledge of the attacked classifier nor explanation method. Overall, our work shows that explanations in text classifiers are very fragile and users need to carefully address their robustness before relying on them in critical applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 被引用 18 次
- Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model PredictionsJingtan Wang, Xiaoqiang Lin, Rui Qiao, Chuan-Sheng Foo 等ICML 2024 · 被引用 12 次
- Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution ProtectionFan Wang, Adams Wai-Kin KongNeurIPS 2022 · 被引用 12 次
- "Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial AttackChristopher Burger, Lingwei Chen, Thai LeEMNLP 2023 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target AttacksBohan Wang, Zewen Liu, Lu Lin, Hui Liu 等ICML 2026 · 被引用 1 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- ColorFool: Semantic Adversarial ColorizationAli Shahin Shamsabadi, Ricardo Sánchez-Matilla, Andrea CavallaroCVPR 2020
- On the Transferability of Adversarial Attacks against Neural Text ClassifierLiping Yuan, Xiaoqing Zheng, Yi Zhou, Cho-Jui Hsieh 等EMNLP 2021 · 被引用 19 次
- Corrupting Neuron Explanations of Deep Visual FeaturesDivyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei WengICCV 2023 · 被引用 3 次
