EDGE: Explaining Deep Reinforcement Learning Policies
Wenbo Guo, Xian Wu, Usmann Khan, Xinyu Xing
摘要
With the rapid development of deep reinforcement learning (DRL) techniques, there is an increasing need to understand and interpret DRL policies. While recent research has developed explanation methods to interpret how an agent determines its moves, they cannot capture the importance of actions/states to a game's final result. In this work, we propose a novel self-explainable model that augments a Gaussian process with a customized kernel function and an interpretable predictor. Together with the proposed model, we also develop a parameter learning procedure that leverages inducing points and variational inference to improve learning efficiency. Using our proposed model, we can predict an agent's final rewards from its game episodes and extract time step importance within episodes as strategy-level explanations for that agent. Through experiments on Atari and MuJoCo games, we verify the explanation fidelity of our method and demonstrate how to employ interpretation to understand agent behavior, discover policy vulnerabilities, remediate policy errors, and even defend against adversarial attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Inherently Explainable Reinforcement Learning in Natural LanguageXiangyu Peng, Mark O. Riedl, Prithviraj AmmanabroluNeurIPS 2022 · 被引用 29 次
- Baffle: Hiding Backdoors in Offline Reinforcement Learning DatasetsChen Gong, Zhou Yang, Yunpeng Bai, Junda He 等S&P 2024 · 被引用 28 次
- PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement LearningJunfeng Guo, Ang Li, Lixu Wang, Cong LiuICCV 2023 · 被引用 27 次
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun 等NeurIPS 2023 · 被引用 24 次
- Learning to Identify Critical States for Reinforcement Learning from VideosHaozhe Liu, Mingchen Zhuge, Bing Li, Yuhui Wang 等ICCV 2023 · 被引用 14 次
它引用的顶会 Paper20
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Invariant RationalizationShiyu Chang, Yang Zhang, Mo Yu, Tommi S. JaakkolaICML 2020 · 被引用 232 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
相关 Paper
- PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning AgentJunfeng Guo, Heng HuangICML 2026
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao 等KDD 2020 · 被引用 27 次
- Self-Interpretable Time Series Prediction with Counterfactual ExplanationsJingquan Yan, Hao WangICML 2023 · 被引用 29 次
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 被引用 28 次
- LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsQianyue Hao, Yiwen Song, Qingmin Liao, Jian Yuan 等NeurIPS 2025 · 被引用 6 次
