Monotonic Robust Policy Optimization with Model Discrepancy
Yuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou, Hongkai Xiong
Abstract
State-of-the-art deep reinforcement learning (DRL) algorithms tend to overfit in some specific environments due to the lack of data diversity in training. To mitigate the model discrepancy between training and target (testing) environments, domain randomization (DR) can generate plenty of environments with a sufficient diversity by randomly sampling environment parameters in simulator. Though standard DR using a uniform distribution improves the average performance on the whole range of environments, the worst-case environment is usually neglected without any performance guarantee. Since the average and worst-case performance are equally important for the generalization in RL, in this paper, we propose a policy optimization approach for concurrently improving the policy's performance in the average case (i.e., over all possible environments) and the worst-case environment. We theoretically derive a lower bound for the worst-case performance of a given policy over all environments. Guided by this lower bound, we formulate an optimization problem which aims to optimize the policy and sampling distribution together, such that the constrained expected performance of all environments is maximized. We prove that the worst-case performance is monotonically improved by iteratively solving this optimization problem. Based on the proposed lower bound, we develop a practical algorithm, named monotonic robust policy optimization (MRPO), and validate MRPO on several robot control tasks. By modifying the environment parameters in simulation, we obtain environments for the same task but with different transition dynamics for training and testing. We demonstrate that MRPO can improve both the average and worst-case performance in the training environments, and facilitate the learned policy with a better generalization capability in unseen testing environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09b6119c-c1ae-4bed-8dc3-ac71bd2463d8Cited by top-tier papers9
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy OptimizationYufei Kuang, Miao Lu, Jie Wang, Qi Zhou et al.AAAI 2022 · 29 citations
- Adjustable Robust Reinforcement Learning for Online 3D Bin PackingYuxin Pan, Yize Chen, Fangzhen LinNeurIPS 2023 · 23 citations
- A Simple Yet Effective Strategy to Robustify the Meta Learning ParadigmQi Wang, Yiqin Lv, Yang-He Feng, Zheng Xie et al.NeurIPS 2023 · 17 citations
- PID-Inspired Inductive Biases for Deep Reinforcement Learning in Partially Observable Control TasksIan Char, Jeff SchneiderNeurIPS 2023 · 8 citations
Builds on2
- Toward A Thousand Lights: Decentralized Deep Reinforcement Learning for Large-Scale Traffic Signal ControlChacha Chen, Hua Wei, Nan Xu, Guanjie Zheng et al.AAAI 2020 · 450 citations
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
Related papers
- Revisiting Domain Randomization via Relaxed State-Adversarial Policy OptimizationYun-Hsuan Lien, Ping-Chun Hsieh, Yu-Shuen WangICML 2023 · 1 citation
- Max-Min Off-Policy Actor-Critic Method Focusing on Worst-Case Robustness to Model MisspecificationTakumi Tanabe, Rei Sato, Kazuto Fukuchi, Jun Sakuma et al.NeurIPS 2022 · 17 citations
- Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceWeiye Zhao, Feihan Li, Yifan Sun, Rui Chen et al.ICML 2024 · 5 citations
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong et al.ICML 2022 · 72 citations
- Domain Randomization via Entropy MaximizationGabriele Tiboni, Pascal Klink, Jan Peters, Tatiana Tommasi et al.ICLR 2024 · 24 citations
