Explainability Via Causal Self-Talk
Nicholas A. Roy, Junkyung Kim, Neil C. Rabinowitz
Abstract
Explaining the behavior of AI systems is an important problem that, in practice, is generally avoided. While the XAI community has been developing an abundance of techniques, most incur a set of costs that the wider deep learning community has been unwilling to pay in most situations. We take a pragmatic view of the issue, and define a set of desiderata that capture both the ambitions of XAI and the practical constraints of deep learning. We describe an effective way to satisfy all the desiderata: train the AI system to build a causal model of itself. We develop an instance of this solution for Deep RL agents: Causal Self-Talk. CST operates by training the agent to communicate with itself across time. We implement this method in a simulated 3D environment, and show how it enables agents to generate faithful and semantically-meaningful explanations of their own behavior. Beyond explanations, we also demonstrate that these learned models provide new ways of building semantic control interfaces to AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1a4885f-e616-42c2-be07-19330f4dc368Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner et al.ICML 2022 · 104 citations
- What Did You Think Would Happen? Explaining Agent Behaviour through Intended OutcomesHerman Yau, Chris Russell, Simon HadfieldNeurIPS 2020 · 44 citations
Related papers
- CausalXRL: Explainable Reinforcement Learning through Causal Graph ReasoningYanming Zhang, Eric Papenhausen, Klaus MuellerICML 2026
- Generating High-Quality Explanations for Navigation in Partially-Revealed EnvironmentsGregory J. SteinNeurIPS 2021 · 19 citations
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 2 citations
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie et al.KDD 2023 · 22 citations
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 28 citations
