Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow
Chen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee, Simon See, Chun-Yi Lee
摘要
Existing Maximum-Entropy (MaxEnt) Reinforcement Learning (RL) methods for continuous action spaces are typically formulated based on actor-critic frameworks and optimized through alternating steps of policy evaluation and policy improvement. In the policy evaluation steps, the critic is updated to capture the soft Q-function. In the policy improvement steps, the actor is adjusted in accordance with the updated soft Q-function. In this paper, we introduce a new MaxEnt RL framework modeled using Energy-Based Normalizing Flows (EBFlow). This framework integrates the policy evaluation steps and the policy improvement steps, resulting in a single objective training process. Our method enables the calculation of the soft value function used in the policy evaluation target without Monte Carlo approximation. Moreover, this design supports the modeling of multi-modal action distributions while facilitating efficient action sampling. To evaluate the performance of our method, we conducted experiments on the MuJoCo benchmark suite and a number of high-dimensional robotic tasks simulated by Omniverse Isaac Gym. The evaluation results demonstrate that our method achieves superior performance compared to widely-adopted representative baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li 等ICLR 2026 · 被引用 41 次
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningShutong Ding, Ke Hu, Shan Zhong, Haoyang Luo 等NeurIPS 2025 · 被引用 22 次
- Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM UnlearningNaixin Zhai, Pengyang Shao, Binbin Zheng, Yonghui Yang 等ACL 2026 · 被引用 10 次
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 被引用 6 次
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy ConstraintsZilin Kang, Chonghua Liao, Tingqiang Xu, Huazhe XuICLR 2026 · 被引用 5 次
它引用的顶会 Paper8
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos 等AAAI 2021 · 被引用 506 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without SamplingWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICML 2020 · 被引用 93 次
- Relative gradient optimization of the Jacobian term in unsupervised deep learningLuigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf 等NeurIPS 2020 · 被引用 25 次
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang 等ICLR 2024 · 被引用 21 次
相关 Paper
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- Extreme Q-Learning: MaxEnt RL without EntropyDivyansh Garg, Joey Hejna, Matthieu Geist, Stefano ErmonICLR 2023 · 被引用 5 次
- Flow-based Domain Randomization for Learning and Sequencing Robotic SkillsAidan Curtis, Eric Li, Michael Noseworthy, Nishad Gothoskar 等ICML 2025
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge MatchingLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun 等ICML 2026 · 被引用 5 次
