Towards Interpretable Deep Reinforcement Learning with Human-Friendly Prototypes
Eoin M. Kenny, Mycal Tucker, Julie Shah
Abstract
Despite recent success of deep learning models in research settings, their application in sensitive domains remains limited because of their opaque decision-making processes. Taking to this challenge, people have proposed various eXplainable AI (XAI) techniques designed to calibrate trust and understandability of black-box models, with the vast majority of work focused on supervised learning. Here, we focus on making an "interpretable-by-design" deep reinforcement learning agent which is forced to use human-friendly prototypes in its decisions, thus making its reasoning process clear. Our proposed method, dubbed Prototype-Wrapper Network (PW-Net), wraps around any neural agent backbone, and results indicate that it does not worsen performance relative to black-box models. Most importantly, we found in a user study that PW-Nets supported better trust calibration and task performance relative to standard interpretability approaches and black-boxes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bf44816c-0357-47c6-ad3d-0d85e5210f96Cited by top-tier papers13
- Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype NetworksQihan Huang, Mengqi Xue, Wenqi Huang, Haofei Zhang et al.ICCV 2023 · 47 citations
- State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User UnderstandingDevleena Das, Sonia Chernova, Been KimNeurIPS 2023 · 33 citations
- Refining Diffusion Planner for Reliable Behavior Synthesis by Automatic Detection of Infeasible PlansKyowoon Lee, Seongun Kim, Jaesik ChoiNeurIPS 2023 · 31 citations
- The Utility of "Even if" Semifactual Explanation to Optimise Positive OutcomesEoin M. Kenny, Weipeng HuangNeurIPS 2023 · 16 citations
- Probabilistic Constrained Reinforcement Learning with Formal InterpretabilityYanran Wang, Qiuchen Qian, David BoyleICML 2024 · 5 citations
Related papers
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 13 citations
- DeXAR: Deep Explainable Sensor-Based Activity Recognition in Smart-Home EnvironmentsLuca Arrotta, Gabriele Civitarese, Claudio BettiniUbiComp 2022 · 40 citations
- Toward Faithful Case-based Reasoning through Learning Prototypes in a Nearest Neighbor-friendly SpaceSeyed Omid Davoudi, Majid KomeiliICLR 2022 · 8 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
