Explainable Reinforcement Learning through a Causal Lens
Prashan Madumal, Tim Miller, Liz Sonenberg, Frank Vetere
摘要
Prominent theories in cognitive science propose that humans understand and represent the knowledge of the world through causal relationships. In making sense of the world, we build causal models in our mind to encode cause-effect relations of events and use these to explain why new events happen by referring to counterfactuals -things that did not happen. In this paper, we use causal models to derive causal explanations of the behaviour of model-free reinforcement learning agents. We present an approach that learns a structural causal model during reinforcement learning and encodes causal relationships between variables of interest. This model is then used to generate explanations of behaviour based on counterfactual analysis of the causal model. We computationally evaluate the model in 6 domains and measure performance and task prediction accuracy. We report on a study with 120 participants who observe agents playing a real-time strategy game (Starcraft II) and then receive explanations of the agents' behaviour. We investigate: 1) participants' understanding gained by explanations through task prediction; 2) explanation satisfaction and 3) trust. Our results show that causal model explanations perform better on these measures compared to two other baseline explanation models. Driven by lack of trust from users and proposed regulations, there are many calls for Artificial Intelligence (AI) systems to become more transparent, interpretable and explainable. This has renewed the interest in Explainable AI (XAI), which has been explored since the expert systems era (Chandrasekaran, Tanner, and Josephson 1989) . A key pillar of XAI is explanation, a justification given for decisions and actions of the system. However, much research and practice in XAI pays little attention to people as intended users of these systems (Miller 2018b). If we are to build systems that are capable of providing 'good' explanations, it is plausible that explanation models should mimic models of human explanation (De Graaf and Malle 2017). Thus, to build
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health CareVenkatesh Sivaraman, Leigh A. Bukowski, Joel Levin, Jeremy M. Kahn 等CHI 2023 · 被引用 126 次
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 被引用 126 次
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 被引用 118 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou 等NeurIPS 2021 · 被引用 78 次
相关 Paper
- CausalXRL: Explainable Reinforcement Learning through Causal Graph ReasoningYanming Zhang, Eric Papenhausen, Klaus MuellerICML 2026
- Explaining Model Confidence Using CounterfactualsThao Le, Tim Miller, Ronal Singh, Liz SonenbergAAAI 2023 · 被引用 9 次
- Explainability Via Causal Self-TalkNicholas A. Roy, Junkyung Kim, Neil C. RabinowitzNeurIPS 2022 · 被引用 10 次
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 被引用 28 次
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan 等ICML 2022 · 被引用 51 次
