Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation
Ruida Zhou, Tao Liu, Min Cheng, Dileep Kalathil, P. R. Kumar, Chao Tian
Abstract
We study robust reinforcement learning (RL) with the goal of determining a well-performing policy that is robust against model mismatch between the training simulator and the testing environment. Previous policy-based robust RL algorithms mainly focus on the tabular setting under uncertainty sets that facilitate robust policy evaluation, but are no longer tractable when the number of states scales up. To this end, we propose two novel uncertainty set formulations, one based on double sampling and the other on an integral probability metric. Both make large-scale robust RL tractable even when one only has access to a simulator. We propose a robust natural actor-critic (RNAC) approach that incorporates the new uncertainty sets and employs function approximation. We provide finite-time convergence guarantees for the proposed RNAC algorithm to the optimal robust policy within the function approximation error. Finally, we demonstrate the robust performance of the policy learned by our proposed RNAC approach in multiple MuJoCo environments and a real-world TurtleBot navigation task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14b9f5f4-3bf6-4af0-a2ae-5c51e85b1faaCited by top-tier papers15
- Provably Fast Convergence of Independent Natural Policy Gradient for Markov Potential GamesYoubang Sun, Tao Liu, Ruida Zhou, P. R. Kumar et al.NeurIPS 2023 · 24 citations
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil et al.NeurIPS 2025 · 18 citations
- Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled PerturbationsYongyuan Liang, Yanchao Sun, Ruijie Zheng, Xiangyu Liu et al.ICLR 2024 · 14 citations
- Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst KernelUri Gadot, Kaixin Wang, Navdeep Kumar, Kfir Yehuda Levy et al.ICML 2024 · 9 citations
- Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity GuaranteesSourav Ganguly, Kishan Panaganti, Arnob Ghosh, Adam WiermanNeurIPS 2025 · 7 citations
Builds on15
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 104 citations
Related papers
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 78 citations
- Robust Reinforcement Learning via Adversarial training with Langevin DynamicsParameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland et al.NeurIPS 2020 · 75 citations
- Doubly Robust Off-Policy Actor-Critic: Convergence and OptimalityTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangICML 2021 · 31 citations
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau et al.ICML 2022 · 25 citations
- Robust Multi-Agent Reinforcement Learning with Model UncertaintyKaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc et al.NeurIPS 2020 · 118 citations
