Maximum Entropy RL (Provably) Solves Some Robust RL Problems
Benjamin Eysenbach, Sergey Levine
Abstract
Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function. In this paper, we prove theoretically that maximum entropy (MaxEnt) RL maximizes a lower bound on a robust RL objective, and thus can be used to learn policies that are robust to some disturbances in the dynamics and the reward function. While this capability of MaxEnt RL has been observed empirically in prior work, to the best of our knowledge our work provides the first rigorous proof and theoretical characterization of the MaxEnt RL robust set. While a number of prior robust RL algorithms have been designed to handle similar disturbances to the reward function or dynamics, these methods typically require additional moving parts and hyperparameters on top of a base RL algorithm. In contrast, our results suggest that MaxEnt RL by itself is robust to certain disturbances, without requiring any additional modifications. While this does not imply that MaxEnt RL is the best available robust RL method, MaxEnt RL is a simple robust RL method with appealing formal guarantees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers66
- Doubly Regularized Markov Decision Processes for Robust Reinforcement LearningYiting He, Zhishuai Liu, Pan XuICML 2026 · 213 citations
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han et al.NeurIPS 2025 · 185 citations
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 104 citations
- Amortizing intractable inference in large language modelsEdward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar et al.ICLR 2024 · 91 citations
- Constrained Decision Transformer for Offline Safe Reinforcement LearningZuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen et al.ICML 2023 · 82 citations
Builds on5
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Robust Reinforcement Learning via Adversarial training with Langevin DynamicsParameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland et al.NeurIPS 2020 · 75 citations
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 68 citations
- Lipschitz Lifelong Reinforcement LearningErwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai et al.AAAI 2021 · 43 citations
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 19 citations
Related papers
- When Maximum Entropy Misleads Policy OptimizationRuipeng Zhang, Ya-Chien Chang, Sicun GaoICML 2025
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
- Robust Inverse Reinforcement Learning under Transition Dynamics MismatchLuca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller et al.NeurIPS 2021 · 40 citations
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICLR 2023 · 9 citations
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau et al.ICML 2022 · 25 citations
