Towards Interpretable Deep Reinforcement Learning with Human-Friendly Prototypes
Eoin M. Kenny, Mycal Tucker, Julie Shah
摘要
Despite recent success of deep learning models in research settings, their application in sensitive domains remains limited because of their opaque decision-making processes. Taking to this challenge, people have proposed various eXplainable AI (XAI) techniques designed to calibrate trust and understandability of black-box models, with the vast majority of work focused on supervised learning. Here, we focus on making an "interpretable-by-design" deep reinforcement learning agent which is forced to use human-friendly prototypes in its decisions, thus making its reasoning process clear. Our proposed method, dubbed Prototype-Wrapper Network (PW-Net), wraps around any neural agent backbone, and results indicate that it does not worsen performance relative to black-box models. Most importantly, we found in a user study that PW-Nets supported better trust calibration and task performance relative to standard interpretability approaches and black-boxes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper13
- Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype NetworksQihan Huang, Mengqi Xue, Wenqi Huang, Haofei Zhang 等ICCV 2023 · 被引用 47 次
- State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User UnderstandingDevleena Das, Sonia Chernova, Been KimNeurIPS 2023 · 被引用 33 次
- Refining Diffusion Planner for Reliable Behavior Synthesis by Automatic Detection of Infeasible PlansKyowoon Lee, Seongun Kim, Jaesik ChoiNeurIPS 2023 · 被引用 31 次
- The Utility of "Even if" Semifactual Explanation to Optimise Positive OutcomesEoin M. Kenny, Weipeng HuangNeurIPS 2023 · 被引用 16 次
- Probabilistic Constrained Reinforcement Learning with Formal InterpretabilityYanran Wang, Qiuchen Qian, David BoyleICML 2024 · 被引用 5 次
相关 Paper
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 被引用 13 次
- DeXAR: Deep Explainable Sensor-Based Activity Recognition in Smart-Home EnvironmentsLuca Arrotta, Gabriele Civitarese, Claudio BettiniUbiComp 2022 · 被引用 40 次
- Toward Faithful Case-based Reasoning through Learning Prototypes in a Nearest Neighbor-friendly SpaceSeyed Omid Davoudi, Majid KomeiliICLR 2022 · 被引用 8 次
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
