A Max-Min Entropy Framework for Reinforcement Learning
Seungyul Han, Youngchul Sung
摘要
In this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the maximum entropy RL guides learning for policies to reach states with high entropy in the future, the proposed max-min entropy framework aims to learn to visit states with low entropy and maximize the entropy of these low-entropy states to promote better exploration. For general Markov decision processes (MDPs), an efficient algorithm is constructed under the proposed max-min entropy framework based on disentanglement of exploration and exploitation. Numerical results show that the proposed algorithm yields drastic performance improvement over the current state-of-the-art RL algorithms. in model-free sample-based learning with function approximation. In order to overcome such limitations associated with implementation of the maximum entropy RL, we propose a max-min entropy framework for RL, which aims to learn policies reaching states with low entropy and maximizing the entropy of these low-entropy states, whereas the conventional maximum entropy RL optimizes for policies that aim to visit states with high entropy and maximize the entropy of those high-entropy states for high entropy of the entire trajectory. We implemented the proposed max-min entropy framework into a practical iterative actor-critic algorithm based on policy iteration with disentangled exploration and exploitation. It is demonstrated that the proposed algorithm significantly enhances exploration capability due to the fairness across states induced by the max-min framework and yields drastic performance improvement over existing RL algorithms including maximum-entropy SAC on difficult control tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Attentive Experience ReplayPeiquan Sun, Wengang Zhou, Houqiang LiAAAI 2020 · 被引用 62 次
- FoX: Formation-Aware Exploration in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Junghyuk Yeom, Seungyul HanAAAI 2024 · 被引用 22 次
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo 等ICML 2024 · 被引用 20 次
- An Adaptive Entropy-Regularization Framework for Multi-Agent Reinforcement LearningWoojun Kim, Youngchul SungICML 2023 · 被引用 20 次
- Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step AlignmentRenye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun 等ICCV 2025 · 被引用 14 次
它引用的顶会 Paper4
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient ExplorationSeungyul Han, Youngchul SungICML 2021 · 被引用 34 次
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 被引用 19 次
相关 Paper
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RLGuojian Zhan, Likun Wang, Pengcheng Wang, Feihong Zhang 等ICML 2026
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang 等ICLR 2024 · 被引用 21 次
- Maximum Entropy Heterogeneous-Agent Reinforcement LearningJiarong Liu, Yifan Zhong, Siyi Hu, Haobo Fu 等ICLR 2024 · 被引用 27 次
- When Maximum Entropy Misleads Policy OptimizationRuipeng Zhang, Ya-Chien Chang, Sicun GaoICML 2025
