Learning to Explain: Generating Stable Explanations Fast
Xuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf, Gholamreza Haffari
Abstract
The importance of explaining the outcome of a machine learning model, especially a blackbox model, is widely acknowledged. Recent approaches explain an outcome by identifying the contributions of input features to this outcome. In environments involving large blackbox models or complex inputs, this leads to computationally demanding algorithms. Further, these algorithms often suffer from low stability, with explanations varying significantly across similar examples. In this paper, we propose a Learning to Explain (L2E) approach that learns the behaviour of an underlying explanation algorithm simultaneously from all training examples. Once the explanation algorithm is distilled into an explainer network, it can be used to explain new instances. Our experiments on three classification tasks, which compare our approach to six explanation algorithms, show that L2E is between 5 and 7.5 × 10 4 times faster than these algorithms, while generating more stable explanations, and having comparable faithfulness to the black-box model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c3249f9-30c9-4db9-b19a-0c523aee8cdaCited by top-tier papers9
- UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionAaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan et al.ICML 2022 · 48 citations
- Tailoring Self-Rationalizers with Multi-Reward DistillationSahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu et al.ICLR 2024 · 23 citations
- Designing a Direct Feedback Loop between Humans and Convolutional Neural Networks through Local ExplanationsTong Steven Sun, Yuyang Gao, Shubham Khaladkar, Sijia Liu et al.CSCW 2023 · 9 citations
- Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language ModelPengwei Zhan, Jing Yang, Xiao Huang, Chunlei Jing et al.ACL 2023 · 1 citation
- Rationalizing Transformer Predictions via End-To-End Differentiable Self-TrainingMarc Felix Brinner, Sina ZarrießEMNLP 2024
Builds on7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong et al.ACL 2020 · 56 citations
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 38 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 4 citations
- Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsSona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam DasSIGMOD 2021 · 1 citation
- LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningXiuyi Jia, Jinchi Li, Yunan Lu, Weiwei LiICML 2025
- An Additive Instance-Wise Approach to Multi-class Model InterpretationVy Vo, Van Nguyen, Trung Le, Quan Hung Tran et al.ICLR 2023
