Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation
Ruida Zhou, Tao Liu, Min Cheng, Dileep Kalathil, P. R. Kumar, Chao Tian
摘要
We study robust reinforcement learning (RL) with the goal of determining a well-performing policy that is robust against model mismatch between the training simulator and the testing environment. Previous policy-based robust RL algorithms mainly focus on the tabular setting under uncertainty sets that facilitate robust policy evaluation, but are no longer tractable when the number of states scales up. To this end, we propose two novel uncertainty set formulations, one based on double sampling and the other on an integral probability metric. Both make large-scale robust RL tractable even when one only has access to a simulator. We propose a robust natural actor-critic (RNAC) approach that incorporates the new uncertainty sets and employs function approximation. We provide finite-time convergence guarantees for the proposed RNAC algorithm to the optimal robust policy within the function approximation error. Finally, we demonstrate the robust performance of the policy learned by our proposed RNAC approach in multiple MuJoCo environments and a real-world TurtleBot navigation task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Provably Fast Convergence of Independent Natural Policy Gradient for Markov Potential GamesYoubang Sun, Tao Liu, Ruida Zhou, P. R. Kumar 等NeurIPS 2023 · 被引用 24 次
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil 等NeurIPS 2025 · 被引用 18 次
- Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled PerturbationsYongyuan Liang, Yanchao Sun, Ruijie Zheng, Xiangyu Liu 等ICLR 2024 · 被引用 14 次
- Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst KernelUri Gadot, Kaixin Wang, Navdeep Kumar, Kfir Yehuda Levy 等ICML 2024 · 被引用 9 次
- Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity GuaranteesSourav Ganguly, Kishan Panaganti, Arnob Ghosh, Adam WiermanNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper15
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 被引用 130 次
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 被引用 104 次
相关 Paper
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 被引用 78 次
- Robust Reinforcement Learning via Adversarial training with Langevin DynamicsParameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland 等NeurIPS 2020 · 被引用 75 次
- Doubly Robust Off-Policy Actor-Critic: Convergence and OptimalityTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangICML 2021 · 被引用 31 次
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau 等ICML 2022 · 被引用 25 次
- Robust Multi-Agent Reinforcement Learning with Model UncertaintyKaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc 等NeurIPS 2020 · 被引用 118 次
