Revisiting Domain Randomization via Relaxed State-Adversarial Policy Optimization
Yun-Hsuan Lien, Ping-Chun Hsieh, Yu-Shuen Wang
摘要
Domain randomization (DR) is widely used in reinforcement learning (RL) to bridge the gap between simulation and reality by maximizing its average returns under the perturbation of environmental parameters. However, even the most complex simulators cannot capture all details in reality due to finite domain parameters and simplified physical models. Additionally, the existing methods often assume that the distribution of domain parameters belongs to a specific family of probability functions, such as normal distributions, which may not be correct. To overcome these limitations, we propose a new approach to DR by rethinking it from the perspective of adversarial state perturbation, without the need for reconfiguring the simulator or relying on prior knowledge about the environment. We also address the issue of over-conservatism that can occur when perturbing agents to the worst states during training by introducing a Relaxed State-Adversarial Algorithm that simultaneously maximizes the average-case and worst-case returns. We evaluate our method by comparing it to state-of-the-art methods, providing experimental results and theoretical proofs to verify its effectiveness. Our source code and appendix are available at https://github.com/sophialien/RAPPO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Robust Deep Reinforcement Learning through Adversarial LossTuomas P. Oikarinen, Wang Zhang, Alexandre Megretski, Luca Daniel 等NeurIPS 2021 · 被引用 134 次
相关 Paper
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou 等ICML 2021 · 被引用 24 次
- Understanding Domain Randomization for Sim-to-real TransferXiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li 等ICLR 2022 · 被引用 164 次
- Domain Randomization via Entropy MaximizationGabriele Tiboni, Pascal Klink, Jan Peters, Tatiana Tommasi 等ICLR 2024 · 被引用 24 次
- Statistical Guarantees for Offline Domain RandomizationArnaud Fickinger, Abderrahim Bendahi, Stuart RussellICLR 2026
- EASI: Evolutionary Adversarial Simulator Identification for Sim-to-Real TransferHaoyu Dong, Huiqiao Fu, Wentao Xu, Zhehao Zhou 等NeurIPS 2024 · 被引用 7 次
