"Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial Attack
Christopher Burger, Lingwei Chen, Thai Le
Abstract
LIME has emerged as one of the most commonly referenced tools in explainable AI (XAI) frameworks that is integrated into critical machine learning applications (e.g., healthcare and finance). However, its stability remains little explored, especially in the context of text data, due to the unique text-space constraints. To address these challenges, in this paper, we first evaluate the inherent instability of LIME on text data to establish a baseline, and then propose a novel algorithm XAIFooler to perturb text inputs and manipulate explanations that casts investigation on the stability of LIME as a text perturbation optimization problem. XAIFooler conforms to the constraints to preserve text semantics and original prediction with small perturbations, and introduces Rank-biased Overlap (RBO) as a key part to guide the optimization of XAIFooler that satisfies all the requirements for explanation similarity measure. Extensive experiments on real-world text datasets demonstrate that XAIFooler significantly outperforms all baselines by large margins in its ability to manipulate LIME’s explanations with high semantic preservability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Fooling Explanations in Text ClassifiersAdam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal FrossardICLR 2022 · 23 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji et al.USENIX Security 2020
Related papers
- S-LIME: Stabilized-LIME for Model ExplanationZhengze Zhou, Giles Hooker, Fei WangKDD 2021 · 98 citations
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 7 citations
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 56 citations
- Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsSona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam DasSIGMOD 2021 · 1 citation
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
