CausalXRL: Explainable Reinforcement Learning through Causal Graph Reasoning
Yanming Zhang, Eric Papenhausen, Klaus Mueller
Abstract
Reinforcement learning is a powerful paradigm for training autonomous agents and has achieved impressive performance in complex environments. However, this success often comes at the cost of interpretability, diminishing trust and complicating efforts to debug and improve agent behavior. To address these challenges, we introduce CausalXRL, a novel framework for explainable reinforcement learning (XRL). A key feature of CausalXRL is its use of causal graph reasoning, which provides transparent, structured, multi-level explanations of agent decision-making. We validate CausalXRL through comprehensive case studies and a two-part evaluation: (1) a quantitative analysis of explanation fidelity and causal-structure learning efficiency in benchmark RL environments, and (2) a qualitative expert study assessing explainability in the real-time strategy (RTS) benchmark MicroRTS. The quantitative results show that CausalXRL can provide faithful explanations while efficiently learning causal structures, and the qualitative expert study suggests that participants found CausalXRL useful for inspecting high-level RTS strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76da5303-6717-4a59-b14c-35302ef920eaBuilds on5
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Causal Influence Detection for Improving Efficiency in Reinforcement LearningMaximilian Seitzer, Bernhard Schölkopf, Georg MartiusNeurIPS 2021 · 120 citations
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer et al.NeurIPS 2024 · 32 citations
- LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement LearningZhuorui Ye, Stephanie Milani, Geoffrey J. Gordon, Fei FangICLR 2025
Related papers
- Inherently Explainable Reinforcement Learning in Natural LanguageXiangyu Peng, Mark O. Riedl, Prithviraj AmmanabroluNeurIPS 2022 · 29 citations
- Explainability Via Causal Self-TalkNicholas A. Roy, Junkyung Kim, Neil C. RabinowitzNeurIPS 2022 · 10 citations
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 59 citations
- UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement LearningShicheng Liu, Minghui ZhuICLR 2025
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 41 citations
