Explainable Reinforcement Learning through a Causal Lens
Prashan Madumal, Tim Miller, Liz Sonenberg, Frank Vetere
Abstract
Prominent theories in cognitive science propose that humans understand and represent the knowledge of the world through causal relationships. In making sense of the world, we build causal models in our mind to encode cause-effect relations of events and use these to explain why new events happen by referring to counterfactuals -things that did not happen. In this paper, we use causal models to derive causal explanations of the behaviour of model-free reinforcement learning agents. We present an approach that learns a structural causal model during reinforcement learning and encodes causal relationships between variables of interest. This model is then used to generate explanations of behaviour based on counterfactual analysis of the causal model. We computationally evaluate the model in 6 domains and measure performance and task prediction accuracy. We report on a study with 120 participants who observe agents playing a real-time strategy game (Starcraft II) and then receive explanations of the agents' behaviour. We investigate: 1) participants' understanding gained by explanations through task prediction; 2) explanation satisfaction and 3) trust. Our results show that causal model explanations perform better on these measures compared to two other baseline explanation models. Driven by lack of trust from users and proposed regulations, there are many calls for Artificial Intelligence (AI) systems to become more transparent, interpretable and explainable. This has renewed the interest in Explainable AI (XAI), which has been explored since the expert systems era (Chandrasekaran, Tanner, and Josephson 1989) . A key pillar of XAI is explanation, a justification given for decisions and actions of the system. However, much research and practice in XAI pays little attention to people as intended users of these systems (Miller 2018b). If we are to build systems that are capable of providing 'good' explanations, it is plausible that explanation models should mimic models of human explanation (De Graaf and Malle 2017). Thus, to build
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b38c0d4-ed10-4922-ab2c-e818ae2b143dCited by top-tier papers37
- Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health CareVenkatesh Sivaraman, Leigh A. Bukowski, Joel Levin, Jeremy M. Kahn et al.CHI 2023 · 126 citations
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 126 citations
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 118 citations
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou et al.NeurIPS 2021 · 78 citations
Related papers
- CausalXRL: Explainable Reinforcement Learning through Causal Graph ReasoningYanming Zhang, Eric Papenhausen, Klaus MuellerICML 2026
- Explaining Model Confidence Using CounterfactualsThao Le, Tim Miller, Ronal Singh, Liz SonenbergAAAI 2023 · 9 citations
- Explainability Via Causal Self-TalkNicholas A. Roy, Junkyung Kim, Neil C. RabinowitzNeurIPS 2022 · 10 citations
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 28 citations
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan et al.ICML 2022 · 51 citations
