Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior Knowledge
Laura Rieger, Chandan Singh, W. James Murdoch, Bin Yu
摘要
For an explanation of a deep learning model to be effective, it must provide both insight into a model and suggest a corresponding action in order to achieve some objective. Too often, the litany of proposed explainable deep learning methods stop at the first step, providing practitioners with insight into a model, but no way to act on it. In this paper, we propose contextual decomposition explanation penalization (CDEP), a method which enables practitioners to leverage existing explanation methods in order to increase the predictive accuracy of deep learning models. In particular, when shown that a model has incorrectly assigned importance to some features, CDEP enables practitioners to correct these errors by directly regularizing the provided explanations. Using explanations provided by contextual decomposition (CD) (Murdoch et al., 2018), we demonstrate the ability of our method to increase performance on an array of toy and real datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu 等NeurIPS 2020 · 被引用 316 次
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 被引用 272 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Model Patching: Closing the Subgroup Performance Gap with Data AugmentationKaran Goel, Albert Gu, Sharon Li, Christopher RéICLR 2021 · 被引用 131 次
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 被引用 102 次
相关 Paper
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 被引用 27 次
- DeciX: Explain Deep Learning Based Code Generation ApplicationsSimin Chen, Zexin Li, Wei Yang, Cong LiuFSE 2024 · 被引用 1 次
- Regional Tree Regularization for Interpretability in Deep Neural NetworksMike Wu, Sonali Parbhoo, Michael C. Hughes, Ryan Kindle 等AAAI 2020 · 被引用 42 次
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen 等AAAI 2021 · 被引用 25 次
- Explainability as statistical inferenceHugo Henri Joseph Senetaire, Damien Garreau, Jes Frellsen, Pierre-Alexandre MatteiICML 2023 · 被引用 4 次
