On the Stability and Convergence of Robust Adversarial Reinforcement Learning: A Case Study on Linear Quadratic Systems
Kaiqing Zhang, Bin Hu, Tamer Basar
摘要
Reinforcement learning (RL) algorithms can fail to generalize due to the gap between the simulation and the real world. One standard remedy is to use robust adversarial RL (RARL) that accounts for this gap during the policy training, by modeling the gap as an adversary against the training agent. In this work, we reexamine the effectiveness of RARL under a fundamental robust control setting: the linear quadratic (LQ) case. We first observe that the popular RARL scheme that greedily alternates agents' updates can easily destabilize the system. Motivated by this, we propose several other policy-based RARL algorithms whose convergence behaviors are then studied both empirically and theoretically. We find: i) the conventional RARL framework (Pinto et al., 2017) can learn a destabilizing policy if the initial policy does not enjoy the robust stability property against the adversary; and ii) with robustly stabilizing initializations, our proposed double-loop RARL algorithm provably converges to the global optimal cost while maintaining robust stability on-the-fly. We also examine the stability and convergence issues of other variants of policy-based RARL algorithms, and then discuss several ways to learn robustly stabilizing initializations. From a robust control perspective, we aim to provide some new and critical angles about RARL, by identifying and addressing the stability issues in this fundamental LQ setting in continuous control. Our results make an initial attempt toward better theoretical understandings of policy-based RARL, the core approach in Pinto et al., 2017.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Robust Reinforcement Learning: A Case Study in Linear Quadratic RegulationBo Pang, Zhong-Ping JiangAAAI 2021 · 被引用 43 次
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 被引用 43 次
- Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian GameSimin Li, Jun Guo, Jingqiao Xiu, Ruixiao Xu 等ICLR 2024 · 被引用 30 次
- Robust Adversarial Reinforcement Learning with Dissipation Inequation ConstraintPeng Zhai, Jie Luo, Zhiyan Dong, Lihua Zhang 等AAAI 2022 · 被引用 29 次
它引用的顶会 Paper1
相关 Paper
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- Zero-Sum Positional Differential Games as a Framework for Robust Reinforcement Learning: Deep Q-Learning ApproachAnton Plaksin, Vitaly KalevICML 2024 · 被引用 1 次
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 被引用 58 次
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 被引用 79 次
- Natural Actor-Critic for Robust Reinforcement Learning with Function ApproximationRuida Zhou, Tao Liu, Min Cheng, Dileep Kalathil 等NeurIPS 2023 · 被引用 55 次
