When Maximum Entropy Misleads Policy Optimization
Ruipeng Zhang, Ya-Chien Chang, Sicun Gao
Abstract
The Maximum Entropy Reinforcement Learning (MaxEnt RL) framework is a leading approach for achieving efficient learning and robust performance across many RL tasks. However, MaxEnt methods have also been shown to struggle with performance-critical control problems in practice, where non-MaxEnt algorithms can successfully learn. In this work, we analyze how the trade-off between robustness and optimality affects the performance of MaxEnt algorithms in complex control tasks: while entropy maximization enhances exploration and robustness, it can also mislead policy optimization, leading to failure in tasks that require precise, low-entropy policies. Through experiments on a variety of control problems, we concretely demonstrate this misleading effect. Our analysis leads to better understanding of how to balance reward design and entropy maximization in challenging control problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb161319-3e68-4c9d-b5ca-b9cd8adf35d5Cited by top-tier papers1
Ask how each one uses itBuilds on6
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
- Understanding and Mitigating the Tradeoff between Robustness and AccuracyAditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi et al.ICML 2020 · 252 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 44 citations
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
Related papers
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 24 citations
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel et al.ICLR 2025
