Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy Improvement
Samuel Neumann, Sungsu Lim, Ajin George Joseph, Yangchen Pan, Adam White, Martha White
摘要
Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the actionvalues, with the addition of entropy regularization for soft variants. In this work, we explore an alternative update for the actor, based on an extension of the cross entropy method (CEM) to condition on inputs (states). The idea is to start with a broader policy and slowly concentrate around maximally valued actions, using a maximum likelihood update towards actions in the top percentile per state. The speed of this concentration is controlled by a proposal policy, that concentrates at a slower rate than the actor. We first provide a policy improvement result in an idealized setting, and then prove that our conditional CEM (CCEM) strategy tracks a CEM update per state, even with changing action-values. We empirically show that our GreedyAC algorithm, that uses CCEM for the actor update, performs better than Soft Actor-Critic and is much less sensitive to entropy-regularization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Value Improved Actor Critic AlgorithmsYaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok 等NeurIPS 2025 · 被引用 7 次
- q-exponential family for policy optimizationLingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai 等ICLR 2025
- Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse PoliciesLingwei Zhu, Han Wang, Yukie NagaiICLR 2025
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
它引用的顶会 Paper6
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 被引用 125 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 被引用 30 次
相关 Paper
- Risk-sensitive control as inference with Rényi divergenceKaito Ito, Kenji KashimaNeurIPS 2024 · 被引用 6 次
- Actor-critic is implicitly biased towards high entropy optimal policiesYuzheng Hu, Ziwei Ji, Matus TelgarskyICLR 2022 · 被引用 12 次
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang 等ICLR 2024 · 被引用 21 次
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient ExplorationSeungyul Han, Youngchul SungICML 2021 · 被引用 34 次
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo 等ICML 2024 · 被引用 20 次
