A Causality Inspired Framework for Model Interpretation
Chenwang Wu, Xiting Wang, Defu Lian, Xing Xie, Enhong Chen
Abstract
A critical issue in eXplainable Artificial Intelligence (XAI) is determining whether explanations uncover the underlying causal factors for model behavior or merely show coincidental relationships. Failing to make this distinction can lead to incorrect understandings. To address this issue, we first understand the model interpretation through a causal lens. We find that the explanation scores of certain representative explanation methods align with the concept of average treatment effect in causal inference and evaluate their relative strengths and limitations from a unified causal perspective. Based on our observations, we outline the major challenges in applying causal inference to model interpretation, including identifying common causes that can be generalized across instances and ensuring that explanations provide a complete causal explanation of model predictions. We then present CIMI, a Causality-Inspired Model Interpreter, which addresses these challenges. CIMI has three modules: the causal sufficiency module and the causal intervention module ensure the explanations are both causally sufficient and generalizable, while the causal prior module facilitates easy learning. Our experiments show that CIMI provides superior and generalizable explanations and is useful for debugging and improving models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 555812aa-578d-437d-9f5c-5df389cc0d63Cited by top-tier papers9
- Uncovering Safety Risks of Large Language Models through Concept Activation VectorZhihao Xu, Ruixuan Huang, Changyu Chen, Xiting WangNeurIPS 2024 · 83 citations
- Customizing Language Models with Instance-wise LoRA for Sequential RecommendationXiaoyu Kong, Jiancan Wu, An Zhang, Leheng Sheng et al.NeurIPS 2024 · 66 citations
- Popularity-Aware Alignment and Contrast for Mitigating Popularity BiasMiaomiao Cai, Lei Chen, Yifan Wang, Haoyue Bai et al.KDD 2024 · 23 citations
- See or Guess: Counterfactually Regularized Image CaptioningQian Cao, Xu Chen, Ruihua Song, Xiting Wang et al.ACM MM 2024 · 7 citations
- Mitigating Distribution Shifts in Sequential Recommendation: An Invariance PerspectiveYuxin Liao, Yonghui Yang, Min Hou, Le Wu et al.SIGIR 2025 · 6 citations
Builds on8
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien et al.NeurIPS 2020 · 295 citations
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang et al.CVPR 2022 · 190 citations
- Reinforcement Subgraph Reasoning for Fake News DetectionRuichao Yang, Xiting Wang, Yiqiao Jin, Chaozhuo Li et al.KDD 2022 · 57 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 2 citations
- Debiasing Concept-based Explanations with Causal AnalysisMohammad Taha Bahadori, David HeckermanICLR 2021 · 8 citations
- Compositional Causal Reasoning Evaluation in Language ModelsJacqueline R. M. A. Maasch, Alihan Hüyük, Xinnuo Xu, Aditya V. Nori et al.ICML 2025
- Theoretical Behavior of XAI Methods in the Presence of Suppressor VariablesRick Wilming, Leo Kieslich, Benedict Clark, Stefan HaufeICML 2023 · 17 citations
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha et al.ICML 2026
