AIRS: Explanation for Deep Reinforcement Learning based Security Applications
Jiahao Yu, Wenbo Guo, Qi Qin, Gang Wang, Ting Wang, Xinyu Xing
摘要
Recently, we have witnessed the success of deep reinforcement learning (DRL) in many security applications, ranging from malware mutation to selfish blockchain mining. Like all other machine learning methods, the lack of explainability has been limiting its broad adoption as users have difficulty establishing trust in DRL models' decisions. Over the past years, different methods have been proposed to explain DRL models but unfortunately, they are often not suitable for security applications, in which explanation fidelity, efficiency, and the capability of model debugging are largely lacking. In this work, we propose AIRS, a general framework to explain deep reinforcement learning-based security applications. Unlike previous works that pinpoint important features to the agent's current action, our explanation is at the step level. It models the relationship between the final reward and the key steps that a DRL agent takes, and thus outputs the steps that are most critical towards the final reward the agent has gathered. Using four representative security-critical applications, we evaluate AIRS from the perspectives of explainability, fidelity, stability, and efficiency. We show that AIRS could outperform alternative explainable DRL methods. We also showcase AIRS's utility, demonstrating that our explanation could facilitate the DRL model's failure offset, help users establish trust in a model decision, and even assist the identification of inappropriate reward designs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with ExplanationZelei Cheng, Xian Wu, Jiahao Yu, Sabrina Yang 等ICML 2024 · 被引用 11 次
- Understanding Individual Agent Importance in Multi-Agent System via Counterfactual ReasoningJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie 等AAAI 2025 · 被引用 11 次
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 被引用 10 次
- SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning SystemsOubo Ma, Yuwen Pu, Linkang Du, Yang Dai 等CCS 2024 · 被引用 6 次
它引用的顶会 Paper21
- MaMaDroid: Detecting Android Malware by Building Markov Chains of Behavioral ModelsEnrico Mariconti, Lucky Onwuzurike, Panagiotis Andriotis, Emiliano De Cristofaro 等NDSS 2017 · 被引用 471 次
- Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning AttacksAvi Schwarzschild, Micah Goldblum, Arjun Gupta, John P. Dickerson 等ICML 2021 · 被引用 207 次
- Effective Program Debloating via Reinforcement LearningKihong Heo, Woosuk Lee, Pardis Pashakhanloo, Mayur NaikCCS 2018 · 被引用 175 次
- Enhancing State-of-the-art Classifiers with API Semantics to Detect Evolved Android MalwareXiaohan Zhang, Yuan Zhang, Ming Zhong, Daizong Ding 等CCS 2020 · 被引用 173 次
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 被引用 121 次
相关 Paper
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun 等NeurIPS 2023 · 被引用 24 次
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao 等KDD 2020 · 被引用 27 次
- CrystalBox: Future-Based Explanations for Input-Driven Deep RL SystemsSagar Patel, Sangeetha Abdu Jyothi, Nina NarodytskaAAAI 2024 · 被引用 1 次
- FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security AnalysisYiling He, Jian Lou, Zhan Qin, Kui RenCCS 2023 · 被引用 6 次
- Explainably Safe Reinforcement LearningSabine Rieder, Stefan Pranger, Debraj Chakraborty, Jan Kretínský 等NeurIPS 2025
