PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent
Junfeng Guo, Heng Huang
摘要
While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention and exploration. In particular, recent work has revealed that RL agents are vulnerable to backdoor attacks, where a victim agent behaves normally under standard conditions but executes malicious actions when a specific trigger is activated. Existing backdoor defenses for RL either require access to the agent’s internal parameters, operate only at the model or trajectory level, or are limited to specific attack types. To ensure the security of RL agents, we propose PolicyGuard, a test-time step-level backdoor defense which leverages Gaussian Process (GP) posterior variance and adapts pseudo trajectories to enable uncertainty computation for individual time step. Besides, we also provide theoretical foundations to explain the efficacy of GP posterior variance. Extensive experiments across seven RL games demonstrate that PolicyGuard achieves state-of-the-art detection performance in most cases, with average AUROC of 0.856 for perturbation-based attacks and 0.859 for adversary-agent attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang 等ICCV 2021 · 被引用 128 次
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 被引用 72 次
- Provable Defense against Backdoor Policies in Reinforcement LearningShubham Kumar Bharti, Xuezhou Zhang, Adish Singla, Jerry ZhuNeurIPS 2022 · 被引用 37 次
- Posterior and Computational Uncertainty in Gaussian ProcessesJonathan Wenger, Geoff Pleiss, Marvin Pförtner, Philipp Hennig 等NeurIPS 2022 · 被引用 31 次
相关 Paper
- PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement LearningJunfeng Guo, Ang Li, Lixu Wang, Cong LiuICCV 2023 · 被引用 27 次
- Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement LearningSanyam Vyas, Alberto Caron, Chris Hicks, Pete Burnap 等AAAI 2026
- BIRD: Generalizable Backdoor Detection and Removal for Deep Reinforcement LearningXuan Chen, Wenbo Guo, Guanhong Tao, Xiangyu Zhang 等NeurIPS 2023 · 被引用 15 次
- Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement LearningEthan Rathbun, Wo Wei Lin, Alina Oprea, Christopher AmatoICLR 2026 · 被引用 5 次
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li 等ICML 2024 · 被引用 4 次
