Adversarial Policy Learning in Two-player Competitive Games
Wenbo Guo, Xian Wu, Sui Huang, Xinyu Xing
摘要
In a two-player deep reinforcement learning task, recent work shows an attacker could learn an adversarial policy that triggers a target agent to perform poorly and even react in an undesired way. However, its efficacy heavily relies upon the zero-sum assumption made in the two-player game. In this work, we propose a new adversarial learning algorithm. It addresses the problem by resetting the optimization goal in the learning process and designing a new surrogate optimization function. Our experiments show that our method significantly improves adversarial agents' exploitability compared with the state-of-art attack. Besides, we also discover that our method could augment an agent with the ability to abuse the target game's unfairness. Finally, we show that agents adversarially retrained against our adversarial agents could obtain stronger adversary-resistance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 被引用 150 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian GameSimin Li, Jun Guo, Jingqiao Xiu, Ruixiao Xu 等ICLR 2024 · 被引用 30 次
- Reward Poisoning Attacks on Offline Multi-Agent Reinforcement LearningYoung Wu, Jeremy McMahan, Xiaojin Zhu, Qiaomin XieAAAI 2023 · 被引用 28 次
- PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement LearningJunfeng Guo, Ang Li, Lixu Wang, Cong LiuICCV 2023 · 被引用 27 次
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningJianwen Sun, Tianwei Zhang, Xiaofei Xie, Lei Ma 等AAAI 2020 · 被引用 141 次
- Spatiotemporally Constrained Action Space Attacks on Deep Reinforcement Learning AgentsXian Yeow Lee, Sambit Ghadai, Kai Liang Tan, Chinmay Hegde 等AAAI 2020 · 被引用 65 次
相关 Paper
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 被引用 19 次
- Adversarial Training Should Be Cast as a Non-Zero-Sum GameAlexander Robey, Fabian Latorre, George J. Pappas, Hamed Hassani 等ICLR 2024 · 被引用 16 次
- Toward Evaluating Robustness of Deep Reinforcement Learning with Continuous ControlTsui-Wei Weng, Krishnamurthy (Dj) Dvijotham, Jonathan Uesato, Kai Xiao 等ICLR 2020 · 被引用 34 次
- Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy EvaluationKosuke Nakanishi, Akihiro Kubo, Yuji Yasui, Shin IshiiICML 2025
- PATROL: Provable Defense against Adversarial Policy in Two-player GamesWenbo Guo, Xian Wu, Lun Wang, Xinyu Xing 等USENIX Security 2023
