Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy Improvement
Samuel Neumann, Sungsu Lim, Ajin George Joseph, Yangchen Pan, Adam White, Martha White
Abstract
Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the actionvalues, with the addition of entropy regularization for soft variants. In this work, we explore an alternative update for the actor, based on an extension of the cross entropy method (CEM) to condition on inputs (states). The idea is to start with a broader policy and slowly concentrate around maximally valued actions, using a maximum likelihood update towards actions in the top percentile per state. The speed of this concentration is controlled by a proposal policy, that concentrates at a slower rate than the actor. We first provide a policy improvement result in an idealized setting, and then prove that our conditional CEM (CCEM) strategy tracks a CEM update per state, even with changing action-values. We empirically show that our GreedyAC algorithm, that uses CCEM for the actor update, performs better than Soft Actor-Critic and is much less sensitive to entropy-regularization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 335beecd-21a0-4040-a2ff-7a41d07a4742Cited by top-tier papers4
- Value Improved Actor Critic AlgorithmsYaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok et al.NeurIPS 2025 · 7 citations
- q-exponential family for policy optimizationLingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai et al.ICLR 2025
- Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse PoliciesLingwei Zhu, Han Wang, Yukie NagaiICLR 2025
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
Builds on6
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 125 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 30 citations
Related papers
- Risk-sensitive control as inference with Rényi divergenceKaito Ito, Kenji KashimaNeurIPS 2024 · 6 citations
- Actor-critic is implicitly biased towards high entropy optimal policiesYuzheng Hu, Ziwei Ji, Matus TelgarskyICLR 2022 · 12 citations
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang et al.ICLR 2024 · 21 citations
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient ExplorationSeungyul Han, Youngchul SungICML 2021 · 34 citations
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo et al.ICML 2024 · 20 citations
