S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor Critic
Safa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang, Bo An, Haipeng Chen, Sanjay Chawla
摘要
Learning expressive stochastic policies instead of deterministic ones has been proposed to achieve better stability, sample complexity, and robustness. Notably, in Maximum Entropy Reinforcement Learning (MaxEnt RL), the policy is modeled as an expressive Energy-Based Model (EBM) over the Q-values. However, this formulation requires the estimation of the entropy of such EBMs, which is an open problem. To address this, previous MaxEnt RL methods either implicitly estimate the entropy, resulting in high computational complexity and variance (SQL), or follow a variational inference procedure that fits simplified actor distributions (e.g., Gaussian) for tractability (SAC). We propose Stein Soft Actor-Critic (SAC), a MaxEnt RL algorithm that learns expressive policies without compromising efficiency. Specifically, SAC uses parameterized Stein Variational Gradient Descent (SVGD) as the underlying policy. We derive a closed-form expression of the entropy of such policies. Our formula is computationally efficient and only depends on first-order derivatives and vector products. Empirical results show that SAC yields more optimal solutions to the MaxEnt objective than SQL and SAC in the multi-goal environment, and outperforms SAC and SQL on the MuJoCo benchmark. Our code is available at: https://github.com/SafaMessaoud/S2AC-Energy-Based-RL-with-Stein-Soft-Actor-Critic
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee 等NeurIPS 2024 · 被引用 29 次
- DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under UncertaintyMingxuan Cui, Duo Zhou, Yuxuan Han, Grani A. Hanasusanto 等ICLR 2026 · 被引用 6 次
- Learning Intractable Multimodal Policies with Reparameterization and Diversity RegularizationZiqi Wang, Jiashun Liu, Ling PanNeurIPS 2025 · 被引用 3 次
- Offline Reinforcement Learning with Generative Trajectory PoliciesXinsong Feng, Leshu Tang, Chenan Wang, Haipeng ChenICML 2026 · 被引用 1 次
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
它引用的顶会 Paper13
- Learning Energy-Based Models by Diffusion Recovery LikelihoodRuiqi Gao, Yang Song, Ben Poole, Ying Nian Wu 等ICLR 2021 · 被引用 144 次
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Learning Energy-Based Model with Variational Auto-Encoder as Amortized SamplerJianwen Xie, Zilong Zheng, Ping LiAAAI 2021 · 被引用 57 次
- A Tale of Two Flows: Cooperative Learning of Langevin Flow and Normalizing Flow Toward Energy-Based ModelJianwen Xie, Yaxuan Zhu, Jun Li, Ping LiICLR 2022 · 被引用 53 次
- Learning Energy-Based Generative Models via Coarse-to-Fine Expanding and SamplingYang Zhao, Jianwen Xie, Ping LiICLR 2021 · 被引用 51 次
相关 Paper
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 被引用 44 次
- Latent State Marginalization as a Low-cost Approach for Improving ExplorationDinghuai Zhang, Aaron C. Courville, Yoshua Bengio, Qinqing Zheng 等ICLR 2023 · 被引用 1 次
- Maximum Entropy Heterogeneous-Agent Reinforcement LearningJiarong Liu, Yifan Zhong, Siyi Hu, Haobo Fu 等ICLR 2024 · 被引用 27 次
- Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RLGuojian Zhan, Likun Wang, Pengcheng Wang, Feihong Zhang 等ICML 2026
