StateMask: Explaining Deep Reinforcement Learning through State Mask
Zelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun, Wenbo Guo, Xinyu Xing
摘要
Despite the promising performance of deep reinforcement learning (DRL) agents in many challenging scenarios, the black-box nature of these agents greatly limits their applications in critical domains. Prior research has proposed several explanation techniques to understand the deep learning-based policies in RL. Most existing methods explain why an agent takes individual actions rather than pinpointing the critical steps to its final reward. To fill this gap, we propose StateMask , a novel method to identify the states most critical to the agent’s final reward. The high-level idea of StateMask is to learn a mask net that blinds a target agent and forces it to take random actions at some steps without compromising the agent’s performance. Through careful design, we can theoretically ensure that the masked agent performs similarly to the original agent. We evaluate StateMask in various popular RL environments and show its superiority over existing explainers in explanation fidelity. We also show that StateMask has better utilities, such as launching adversarial attacks and patching policy errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- Understanding Individual Agent Importance in Multi-Agent System via Counterfactual ReasoningJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie 等AAAI 2025 · 被引用 11 次
- Soft-Label Integration for Robust Toxicity ClassificationZelei Cheng, Xian Wu, Jiahao Yu, Shuo Han 等NeurIPS 2024 · 被引用 7 次
- SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksYongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan 等AAAI 2025 · 被引用 4 次
- Explainable Reinforcement Learning from Human Feedback to Improve AlignmentShicheng Liu, Siyuan Xu, Wenjie Qiu, Hangfan Zhang 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper10
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 被引用 108 次
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha 等ICLR 2020 · 被引用 99 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- PerfectDou: Dominating DouDizhu with Perfect Information DistillationGuan Yang, Minghuan Liu, Weijun Hong, Weinan Zhang 等NeurIPS 2022 · 被引用 41 次
相关 Paper
- AIRS: Explanation for Deep Reinforcement Learning based Security ApplicationsJiahao Yu, Wenbo Guo, Qi Qin, Gang Wang 等USENIX Security 2023
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li 等ICML 2024 · 被引用 4 次
- RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with ExplanationZelei Cheng, Xian Wu, Jiahao Yu, Sabrina Yang 等ICML 2024 · 被引用 11 次
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 被引用 13 次
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningJianwen Sun, Tianwei Zhang, Xiaofei Xie, Lei Ma 等AAAI 2020 · 被引用 141 次
