When Maximum Entropy Misleads Policy Optimization
Ruipeng Zhang, Ya-Chien Chang, Sicun Gao
摘要
The Maximum Entropy Reinforcement Learning (MaxEnt RL) framework is a leading approach for achieving efficient learning and robust performance across many RL tasks. However, MaxEnt methods have also been shown to struggle with performance-critical control problems in practice, where non-MaxEnt algorithms can successfully learn. In this work, we analyze how the trade-off between robustness and optimality affects the performance of MaxEnt algorithms in complex control tasks: while entropy maximization enhances exploration and robustness, it can also mislead policy optimization, leading to failure in tasks that require precise, low-entropy policies. Through experiments on a variety of control problems, we concretely demonstrate this misleading effect. Our analysis leads to better understanding of how to balance reward design and entropy maximization in challenging control problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov 等NeurIPS 2020 · 被引用 336 次
- Understanding and Mitigating the Tradeoff between Robustness and AccuracyAditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi 等ICML 2020 · 被引用 252 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 被引用 44 次
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee 等NeurIPS 2024 · 被引用 29 次
相关 Paper
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 被引用 24 次
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki 等ICLR 2020 · 被引用 138 次
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel 等ICLR 2025
