What Did You Think Would Happen? Explaining Agent Behaviour through Intended Outcomes
Herman Yau, Chris Russell, Simon Hadfield
Abstract
We present a novel form of explanation for Reinforcement Learning, based around the notion of intended outcome. These explanations describe the outcome an agent is trying to achieve by its actions. We provide a simple proof that general methods for post-hoc explanations of this nature are impossible in traditional reinforcement learning. Rather, the information needed for the explanations must be collected in conjunction with training the agent. We derive approaches designed to extract local explanations based on intention for several variants of Q-function approximation and prove consistency between the explanations and the Q-values learned. We demonstrate our method on multiple reinforcement learning problems, and provide code 1 to help researchers introspecting their RL environments and algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a999fa29-2ed6-4353-9ce4-bc99fc155125Cited by top-tier papers9
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 30 citations
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 18 citations
- Inverse Contextual Bandits: Learning How Behavior Evolves over TimeAlihan Hüyük, Daniel Jarrett, Mihaela van der SchaarICML 2022 · 14 citations
- Explainability Via Causal Self-TalkNicholas A. Roy, Junkyung Kim, Neil C. RabinowitzNeurIPS 2022 · 10 citations
- Contextualized Policy Recovery: Modeling and Interpreting Medical Decisions with Adaptive Imitation LearningJannik Deuschel, Caleb Ellington, Yingtao Luo, Benjamin J. Lengerich et al.ICML 2024 · 5 citations
Related papers
- Explaining Reinforcement Learning Agents through Counterfactual Action OutcomesYotam Amitai, Yael Septon, Ofra AmirAAAI 2024 · 19 citations
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus et al.NeurIPS 2022 · 18 citations
- Local Explanations for Reinforcement LearningRonny Luss, Amit Dhurandhar, Miao LiuAAAI 2023 · 5 citations
- Outcome-Driven Reinforcement Learning via Variational InferenceTim G. J. Rudner, Vitchyr Pong, Rowan McAllister, Yarin Gal et al.NeurIPS 2021 · 24 citations
- Imitating Past Successes can be Very SuboptimalBenjamin Eysenbach, Soumith Udatha, Russ Salakhutdinov, Sergey LevineNeurIPS 2022 · 28 citations
