Local Explanations for Reinforcement Learning
Ronny Luss, Amit Dhurandhar, Miao Liu
摘要
Many works in explainable AI have focused on explaining black-box classification models. Explaining deep reinforcement learning (RL) policies in a manner that could be understood by domain users has received much less attention. In this paper, we propose a novel perspective to understanding RL policies based on identifying important states from automatically learned meta-states. The key conceptual difference between our approach and many previous ones is that we form meta-states based on locality governed by the expert policy dynamics rather than based on similarity of actions, and that we do not assume any particular knowledge of the underlying topology of the state space. Theoretically, we show that our algorithm to find meta-states converges and the objective that selects important states from each meta-state is submodular leading to efficient high quality greedy selection. Experiments on four domains (four rooms, door-key, minipacman, and pong) and a carefully conducted user study illustrate that our perspective leads to better understanding of the policy. We conjecture that this is a result of our meta-states being more intuitive in that the corresponding important states are strong indicators of tractable intermediate goals that are easier for humans to interpret and follow.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- What Did You Think Would Happen? Explaining Agent Behaviour through Intended OutcomesHerman Yau, Chris Russell, Simon HadfieldNeurIPS 2020 · 被引用 44 次
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning AgentsChristian Rupprecht, Cyril Ibrahim, Christopher J. PalICLR 2020 · 被引用 36 次
- Re-understanding Finite-State Representations of Recurrent Policy NetworksMohamad H. Danesh, Anurag Koul, Alan Fern, Saeed KhorramICML 2021 · 被引用 26 次
相关 Paper
- TripleTree: A Versatile Interpretable Representation of Black Box Agents and their EnvironmentsTom Bewley, Jonathan LawryAAAI 2021 · 被引用 33 次
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus 等NeurIPS 2022 · 被引用 18 次
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 被引用 13 次
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda 等AAAI 2020 · 被引用 37 次
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 被引用 28 次
