Toward Subspace-Perturbed Trajectory-Aware Backdoor Attacks in Deep Reinforcement Learning
Yaguan Qian, Taining Zhang, Qiqi Bao, Yanru Guo, Lufang Zhang, Zhaoquan Gu, Shouling Ji, Bin Wang, Zhen Lei
摘要
Deep Reinforcement Learning agents are in- creasingly used in safety-critical domains but remain vulnerable to stealthy backdoor attacks. Existing outer-loop attacks face a trade-off be- tween perceptual stealth, poisoning efficiency, and value-function consistency, often making the at- tack ineffective or easily exposed. To address these challenges, we propose SpecDRL, a uni- fied framework that ❶ embeds triggers in the least sensitive subspaces of the state manifold via Subspace-Aware Injection, exploiting percep- tual blind spots, ❷ selects the most influential time steps for poisoning through Value-Guided Strategic Sampling based on Return-to-Go and Temporal-Difference error, and ❸ preserves re- ward integrity via Bellman-Consistent Dynamic Reward Poisoning, which analytically enforces ϵ- consistency of value functions and bounds global return deviations. Experiments across 12 Atari en- vironments demonstrate that SpecDRL achieves near-100% attack success, accelerates backdoor convergence, and maintains benign task perfor- mance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning AgentsEthan Rathbun, Christopher Amato, Alina OpreaNeurIPS 2024 · 被引用 27 次
- Efficient Backdoor Attacks for Deep Neural Networks in Real-world ScenariosZiqiang Li, Hong Sun, Pengfei Xia, Heng Li 等ICLR 2024 · 被引用 11 次
- Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement LearningShijie Liu, Andrew C. Cullen, Paul MONTAGUE, Sarah Erfani 等ICML 2026 · 被引用 5 次
相关 Paper
- BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningJing Cui, Yufei Han, Yuzhe Ma, Jianbin Jiao 等AAAI 2024 · 被引用 31 次
- Adversarial Inception Backdoor Attacks against Reinforcement LearningEthan Rathbun, Alina Oprea, Christopher AmatoICML 2025
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li 等ICML 2024 · 被引用 4 次
- Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement LearningSanyam Vyas, Alberto Caron, Chris Hicks, Pete Burnap 等AAAI 2026
- Provable Defense against Backdoor Policies in Reinforcement LearningShubham Kumar Bharti, Xuezhou Zhang, Adish Singla, Jerry ZhuNeurIPS 2022 · 被引用 37 次
