What Did You Think Would Happen? Explaining Agent Behaviour through Intended Outcomes
Herman Yau, Chris Russell, Simon Hadfield
摘要
We present a novel form of explanation for Reinforcement Learning, based around the notion of intended outcome. These explanations describe the outcome an agent is trying to achieve by its actions. We provide a simple proof that general methods for post-hoc explanations of this nature are impossible in traditional reinforcement learning. Rather, the information needed for the explanations must be collected in conjunction with training the agent. We derive approaches designed to extract local explanations based on intention for several variants of Q-function approximation and prove consistency between the explanations and the Q-values learned. We demonstrate our method on multiple reinforcement learning problems, and provide code 1 to help researchers introspecting their RL environments and algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 被引用 30 次
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 被引用 18 次
- Inverse Contextual Bandits: Learning How Behavior Evolves over TimeAlihan Hüyük, Daniel Jarrett, Mihaela van der SchaarICML 2022 · 被引用 14 次
- Explainability Via Causal Self-TalkNicholas A. Roy, Junkyung Kim, Neil C. RabinowitzNeurIPS 2022 · 被引用 10 次
- Contextualized Policy Recovery: Modeling and Interpreting Medical Decisions with Adaptive Imitation LearningJannik Deuschel, Caleb Ellington, Yingtao Luo, Benjamin J. Lengerich 等ICML 2024 · 被引用 5 次
相关 Paper
- Explaining Reinforcement Learning Agents through Counterfactual Action OutcomesYotam Amitai, Yael Septon, Ofra AmirAAAI 2024 · 被引用 19 次
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus 等NeurIPS 2022 · 被引用 18 次
- Local Explanations for Reinforcement LearningRonny Luss, Amit Dhurandhar, Miao LiuAAAI 2023 · 被引用 5 次
- Outcome-Driven Reinforcement Learning via Variational InferenceTim G. J. Rudner, Vitchyr Pong, Rowan McAllister, Yarin Gal 等NeurIPS 2021 · 被引用 24 次
- Imitating Past Successes can be Very SuboptimalBenjamin Eysenbach, Soumith Udatha, Russ Salakhutdinov, Sergey LevineNeurIPS 2022 · 被引用 28 次
