Robust Reinforcement Learning via Adversarial training with Langevin Dynamics
Parameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland, Cheng Shi, Volkan Cevher
Abstract
We introduce a sampling perspective to tackle the challenging task of training robust Reinforcement Learning (RL) agents. Leveraging the powerful Stochastic Gradient Langevin Dynamics, we present a novel, scalable two-player RL algorithm, which is a sampling variant of the two-player policy gradient method. Our algorithm consistently outperforms existing baselines, in terms of generalization across different training and testing conditions, on several MuJoCo environments. Our experiments also show that, even for objective functions that entirely ignore potential environmental shifts, our sampling approach remains highly robust in comparison to standard RL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4faac751-028f-426a-9751-5cc549893b86Cited by top-tier papers25
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- RORL: Robust Offline Reinforcement Learning via Conservative SmoothingRui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang et al.NeurIPS 2022 · 118 citations
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 96 citations
- Policy Smoothing for Provably Robust Reinforcement LearningAounon Kumar, Alexander Levine, Soheil FeiziICLR 2022 · 62 citations
- Robust Predictable ControlBen Eysenbach, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2021 · 53 citations
Related papers
- Natural Actor-Critic for Robust Reinforcement Learning with Function ApproximationRuida Zhou, Tao Liu, Min Cheng, Dileep Kalathil et al.NeurIPS 2023 · 55 citations
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou et al.ICML 2021 · 24 citations
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 27 citations
- Langevin Policy for Safe Reinforcement LearningFenghao Lei, Long Yang, Shiting Wen, Zhixiong Huang et al.ICML 2024 · 2 citations
- Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled PerturbationsYongyuan Liang, Yanchao Sun, Ruijie Zheng, Xiangyu Liu et al.ICLR 2024 · 14 citations
