Optimal Attack and Defense for Reinforcement Learning
Jeremy McMahan, Young Wu, Xiaojin Zhu, Qiaomin Xie
摘要
To ensure the usefulness of Reinforcement Learning (RL) in real systems, it is crucial to ensure they are robust to noise and adversarial attacks. In adversarial RL, an external attacker has the power to manipulate the victim agent's interaction with the environment. We study the full class of online manipulation attacks, which include (i) state attacks, (ii) observation attacks (which are a generalization of perceived-state attacks), (iii) action attacks, and (iv) reward attacks. We show the attacker's problem of designing a stealthy attack that maximizes its own expected reward, which often corresponds to minimizing the victim's value, is captured by a Markov Decision Process (MDP) that we call a meta-MDP since it is not the true environment but a higher level environment induced by the attacked interaction. We show that the attacker can derive optimal attacks by planning in polynomial time or learning with polynomial sample complexity using standard RL techniques. We argue that the optimal defense policy for the victim can be computed as the solution to a stochastic Stackelberg game, which can be further simplified into a partially-observable turn-based stochastic game (POTBSG). Neither the attacker nor the victim would benefit from deviating from their respective optimal policies, thus such solutions are truly robust. Although the defense problem is NP-hard, we show that optimal Markovian defenses can be computed (learned) in polynomial time (sample complexity) in many scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Adversarially Robust Decision TransformerXiaohang Tang, Afonso Marques, Parameswaran Kamalaruban, Ilija BogunovicNeurIPS 2024 · 被引用 5 次
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 被引用 1 次
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement LearningTairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye 等CVPR 2026
- When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPsJose Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao 等ICML 2026
- On Minimizing Adversarial Counterfactual Error in Adversarial Reinforcement LearningRoman Belaire, Arunesh Sinha, Pradeep VarakanthamICLR 2025
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RLYanchao Sun, Ruijie Zheng, Yongyuan Liang, Furong HuangICLR 2022 · 被引用 82 次
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional SmoothingFan Wu, Linyi Li, Zijian Huang, Yevgeniy Vorobeychik 等ICLR 2022 · 被引用 64 次
相关 Paper
- Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement LearningAmin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 等ICML 2020 · 被引用 145 次
- Adversarial Cheap TalkChris Lu, Timon Willi, Alistair Letcher, Jakob Nicolaus FoersterICML 2023 · 被引用 17 次
- Efficient Adversarial Attacks on Online Multi-agent Reinforcement LearningGuanlin Liu, Lifeng LaiNeurIPS 2023 · 被引用 24 次
- Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RLXiangyu Liu, Souradip Chakraborty, Yanchao Sun, Furong HuangICLR 2024 · 被引用 10 次
- Adaptive Reward-Poisoning Attacks against Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish Singla, Xiaojin ZhuICML 2020 · 被引用 154 次
