Malicious Attacks against Deep Reinforcement Learning Interpretations
Mengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao, Aidong Zhang
Abstract
The past years have witnessed the rapid development of deep reinforcement learning (DRL), which is a combination of deep learning and reinforcement learning (RL). However, the adoption of deep neural networks makes the decision-making process of DRL opaque and lacking transparency. Motivated by this, various interpretation methods for DRL have been proposed. However, those interpretation methods make an implicit assumption that they are performed in a reliable and secure environment. In practice, sequential agent-environment interactions expose the DRL algorithms and their corresponding downstream interpretations to extra adversarial risk. In spite of the prevalence of malicious attacks, there is no existing work studying the possibility and feasibility of malicious attacks against DRL interpretations. To bridge this gap, in this paper, we investigate the vulnerability of DRL interpretation methods. Specifically, we introduce the first study of the adversarial attacks against DRL interpretations, and propose an optimization framework based on which the optimal adversarial attack strategy can be derived. In addition, we study the vulnerability of DRL interpretation methods to the model poisoning attacks, and present an algorithmic framework to rigorously formulate the proposed model poisoning attack. Finally, we conduct both theoretical analysis and extensive experiments to validate the effectiveness of the proposed malicious attacks against DRL interpretations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9defd754-c87c-4de0-807c-3db8823c06ceCited by top-tier papers8
- Towards Automating Model Explanations with Certified Robustness GuaranteesMengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao et al.AAAI 2022 · 16 citations
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao et al.ICML 2024 · 10 citations
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 5 citations
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li et al.ICML 2024 · 4 citations
- Belief-Enriched Pessimistic Q-Learning against Adversarial State PerturbationsXiaolin Sun, Zizhan ZhengICLR 2024 · 4 citations
Builds on4
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningJianwen Sun, Tianwei Zhang, Xiaofei Xie, Lei Ma et al.AAAI 2020 · 141 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
- Towards Interpretation of Pairwise LearningMengdi Huai, Di Wang, Chenglin Miao, Aidong ZhangAAAI 2020 · 8 citations
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji et al.USENIX Security 2020
Related papers
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 72 citations
- When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPsJose Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao et al.ICML 2026
- Efficient Adversarial Attacks on Online Multi-agent Reinforcement LearningGuanlin Liu, Lifeng LaiNeurIPS 2023 · 24 citations
- Adversarial Inception Backdoor Attacks against Reinforcement LearningEthan Rathbun, Alina Oprea, Christopher AmatoICML 2025
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 19 citations
