Learning to Explain: Generating Stable Explanations Fast
Xuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf, Gholamreza Haffari
摘要
The importance of explaining the outcome of a machine learning model, especially a blackbox model, is widely acknowledged. Recent approaches explain an outcome by identifying the contributions of input features to this outcome. In environments involving large blackbox models or complex inputs, this leads to computationally demanding algorithms. Further, these algorithms often suffer from low stability, with explanations varying significantly across similar examples. In this paper, we propose a Learning to Explain (L2E) approach that learns the behaviour of an underlying explanation algorithm simultaneously from all training examples. Once the explanation algorithm is distilled into an explainer network, it can be used to explain new instances. Our experiments on three classification tasks, which compare our approach to six explanation algorithms, show that L2E is between 5 and 7.5 × 10 4 times faster than these algorithms, while generating more stable explanations, and having comparable faithfulness to the black-box model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionAaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 等ICML 2022 · 被引用 48 次
- Tailoring Self-Rationalizers with Multi-Reward DistillationSahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu 等ICLR 2024 · 被引用 23 次
- Designing a Direct Feedback Loop between Humans and Convolutional Neural Networks through Local ExplanationsTong Steven Sun, Yuyang Gao, Shubham Khaladkar, Sijia Liu 等CSCW 2023 · 被引用 9 次
- Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language ModelPengwei Zhan, Jing Yang, Xiao Huang, Chunlei Jing 等ACL 2023 · 被引用 1 次
- Rationalizing Transformer Predictions via End-To-End Differentiable Self-TrainingMarc Felix Brinner, Sina ZarrießEMNLP 2024
它引用的顶会 Paper7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong 等ACL 2020 · 被引用 56 次
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 被引用 38 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
相关 Paper
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 被引用 4 次
- Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsSona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam DasSIGMOD 2021 · 被引用 1 次
- LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningXiuyi Jia, Jinchi Li, Yunan Lu, Weiwei LiICML 2025
- An Additive Instance-Wise Approach to Multi-class Model InterpretationVy Vo, Van Nguyen, Trung Le, Quan Hung Tran 等ICLR 2023
