On the Relationship Between Explanation and Prediction: A Causal View
Amir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf, Been Kim
摘要
Being able to provide explanations for a model's decision has become a central requirement for the development, deployment, and adoption of machine learning models. However, we are yet to understand what explanation methods can and cannot do. How do upstream factors such as data, model prediction, hyperparameters, and random initialization influence downstream explanations? While previous work raised concerns that explanations (E) may have little relationship with the prediction (Y), there is a lack of conclusive study to quantify this relationship. Our work borrows tools from causal inference to systematically assay this relationship. More specifically, we study the relationship between E and Y by measuring the treatment effect when intervening on their causal ancestors, i.e., on hyperparameters and inputs used to generate saliency-based Es or Ys. Our results suggest that the relationships between E and Y is far from ideal. In fact, the gap between 'ideal' case only increase in higher-performing models -- models that are likely to be deployed. Our work is a promising first step towards providing a quantitative measure of the relationship between E and Y, which could also inform the future development of methods for E with a quantitative metric.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How Interpretable Are Interpretable Graph Neural Networks?Yongqiang Chen, Yatao Bian, Bo Han, James ChengICML 2024 · 被引用 17 次
- Prospector Heads: Generalized Feature Attribution for Large Models & DataGautam Machiraju, Alexander Derry, Arjun D. Desai, Neel Guha 等ICML 2024 · 被引用 2 次
- Explanations are a Means to an End: Decision Theoretic Explanation EvaluationZiyang Guo, Berk Ustun, Jessica HullmanICML 2026
- A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model InterpretabilityPouria Fatemi, Ehsan Sharifian, Mohammad Hossein YassaeeICML 2025
它引用的顶会 Paper9
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 被引用 249 次
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 被引用 102 次
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 被引用 46 次
相关 Paper
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie 等KDD 2023 · 被引用 22 次
- DeciX: Explain Deep Learning Based Code Generation ApplicationsSimin Chen, Zexin Li, Wei Yang, Cong LiuFSE 2024 · 被引用 1 次
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 被引用 16 次
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 被引用 116 次
