Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions
Zhengxian Lin, Kin-Ho Lam, Alan Fern
Abstract
We investigate a deep reinforcement learning (RL) architecture that supports explaining why a learned agent prefers one action over another. The key idea is to learn action-values that are directly represented via human-understandable properties of expected futures. This is realized via the embedded self-prediction (ESP) model, which learns said properties in terms of human provided features. Action preferences can then be explained by contrasting the future properties predicted for each action. To address cases where there are a large number of features, we develop a novel method for computing minimal sufficient explanations from an ESP. Our case studies in three domains, including a complex strategy game, show that ESP models can be effectively learned and support insightful explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable RepresentationsSarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava et al.ICLR 2022 · 39 citations
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun et al.NeurIPS 2023 · 24 citations
- Explaining Reinforcement Learning Agents through Counterfactual Action OutcomesYotam Amitai, Yael Septon, Ofra AmirAAAI 2024 · 19 citations
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus et al.NeurIPS 2022 · 18 citations
Related papers
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 41 citations
- CausalXRL: Explainable Reinforcement Learning through Causal Graph ReasoningYanming Zhang, Eric Papenhausen, Klaus MuellerICML 2026
- Self-explaining deep models with logic rule reasoningSeungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi et al.NeurIPS 2022 · 27 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Generating High-Quality Explanations for Navigation in Partially-Revealed EnvironmentsGregory J. SteinNeurIPS 2021 · 19 citations
