Adversarial Inception Backdoor Attacks against Reinforcement Learning
Ethan Rathbun, Alina Oprea, Christopher Amato
Abstract
Recent works have demonstrated the vulnerability of Deep Reinforcement Learning (DRL) algorithms against training-time, backdoor poisoning attacks. The objectives of these attacks are twofold: induce pre-determined, adversarial behavior in the agent upon observing a fixed trigger during deployment while allowing the agent to solve its intended task during training. Prior attacks assume arbitrary control over the agent's rewards, inducing values far outside the environment's natural constraints. This results in brittle attacks that fail once the proper reward constraints are enforced. Thus, in this work we propose a new class of backdoor attacks against DRL which are the first to achieve state of the art performance under strict reward constraints. These "inception" attacks manipulate the agent's training data -inserting the trigger into prior observations and replacing high return actions with those of the targeted adversarial behavior. We formally define these attacks and prove they achieve both adversarial objectives against arbitrary Markov Decision Processes (MDP). Using this framework we devise an online inception attack which achieves an 100% attack success rate on multiple environments under constrained rewards while minimally impacting the agent's task performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement LearningEthan Rathbun, Wo Wei Lin, Alina Oprea, Christopher AmatoICLR 2026 · 5 citations
- Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement LearningOubo Ma, Ruixiao Lin, Yang Dai, Jiahao Chen et al.ICML 2026 · 1 citation
- PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning AgentJunfeng Guo, Heng HuangICML 2026
- Toward Subspace-Perturbed Trajectory-Aware Backdoor Attacks in Deep Reinforcement LearningYaguan Qian, Taining Zhang, Qiqi Bao, Yanru Guo et al.ICML 2026
Builds on8
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant et al.ICLR 2020 · 415 citations
- Provable Defense against Backdoor Policies in Reinforcement LearningShubham Kumar Bharti, Xuezhou Zhang, Adish Singla, Jerry ZhuNeurIPS 2022 · 37 citations
- SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning AgentsEthan Rathbun, Christopher Amato, Alina OpreaNeurIPS 2024 · 27 citations
- Optimal Attack and Defense for Reinforcement LearningJeremy McMahan, Young Wu, Xiaojin Zhu, Qiaomin XieAAAI 2024 · 25 citations
- Adversarial Cheap TalkChris Lu, Timon Willi, Alistair Letcher, Jakob Nicolaus FoersterICML 2023 · 17 citations
Related papers
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 72 citations
- Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement LearningSanyam Vyas, Alberto Caron, Chris Hicks, Pete Burnap et al.AAAI 2026
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li et al.ICML 2024 · 4 citations
- TrojanTO: Action-Level Backdoor Attacks Against Trajectory Optimization ModelsYang Dai, Oubo Ma, Xingxing Liang, Longfei Zhang et al.ICLR 2026 · 3 citations
- BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningJing Cui, Yufei Han, Yuzhe Ma, Jianbin Jiao et al.AAAI 2024 · 31 citations
