BadRL: Sparse Targeted Backdoor Attack against Reinforcement Learning
Jing Cui, Yufei Han, Yuzhe Ma, Jianbin Jiao, Junge Zhang
摘要
Backdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work, we propose a novel approach, BadRL, which focuses on conducting highly sparse backdoor poisoning efforts during training and testing while maintaining successful attacks. Our algorithm, BadRL, strategically chooses state observations with high attack values to inject triggers during training and testing, thereby reducing the chances of detection. In contrast to the previous methods that utilize sample-agnostic trigger patterns, BadRL dynamically generates distinct trigger patterns based on targeted state observations, thereby enhancing its effectiveness. Theoretical analysis shows that the targeted backdoor attack is always viable and remains stealthy under specific assumptions. Empirical results on various classic RL tasks illustrate that BadRL can substantially degrade the performance of a victim agent with minimal poisoning efforts (0.003% of total training steps) during training and infrequent attacks during testing. Code is available at: https://github.com/7777777cc/code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等NeurIPS 2024 · 被引用 195 次
- SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning AgentsEthan Rathbun, Christopher Amato, Alina OpreaNeurIPS 2024 · 被引用 27 次
- Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous DrivingXuan Chen, Shiwei Feng, Zikang Xiong, Shengwei An 等NeurIPS 2025 · 被引用 6 次
- TrojanTO: Action-Level Backdoor Attacks Against Trajectory Optimization ModelsYang Dai, Oubo Ma, Xingxing Liang, Longfei Zhang 等ICLR 2026 · 被引用 3 次
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 被引用 1 次
它引用的顶会 Paper14
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 被引用 213 次
相关 Paper
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 被引用 72 次
- Toward Subspace-Perturbed Trajectory-Aware Backdoor Attacks in Deep Reinforcement LearningYaguan Qian, Taining Zhang, Qiqi Bao, Yanru Guo 等ICML 2026
- Adversarial Inception Backdoor Attacks against Reinforcement LearningEthan Rathbun, Alina Oprea, Christopher AmatoICML 2025
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li 等ICML 2024 · 被引用 4 次
- BIRD: Generalizable Backdoor Detection and Removal for Deep Reinforcement LearningXuan Chen, Wenbo Guo, Guanhong Tao, Xiangyu Zhang 等NeurIPS 2023 · 被引用 15 次
