Characterizing the Gap Between Actor-Critic and Policy Gradient
Junfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale Schuurmans
Abstract
Actor-critic (AC) methods are ubiquitous in reinforcement learning. Although it is understood that AC methods are closely related to policy gradient (PG), their precise connection has not been fully characterized previously. In this paper, we explain the gap between AC and PG methods by identifying the exact adjustment to the AC objective/gradient that recovers the true policy gradient of the cumulative reward objective (PG). Furthermore, by viewing the AC method as a two-player Stackelberg game between the actor and critic, we show that the Stackelberg policy gradient can be recovered as a special case of our more general analysis. Based on these results, we develop practical algorithms, Residual Actor-Critic and Stackelberg Actor-Critic, for estimating the correction between AC and PG and use these to modify the standard AC algorithm. Experiments on popular tabular and continuous environments show the proposed corrections can improve both the sample efficiency and final performance of existing AC methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee1cbff9-24d2-4974-ac84-6dd93e7d13e9Cited by top-tier papers7
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 176 citations
- Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning AlgorithmsLiyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov et al.AAAI 2022 · 50 citations
- Exploring Gradient Explosion in Generative Adversarial Imitation Learning: A Probabilistic PerspectiveWanying Wang, Yichen Zhu, Yirui Zhou, Chaomin Shen et al.AAAI 2024 · 13 citations
- A Connection between One-Step RL and Critic Regularization in Reinforcement LearningBenjamin Eysenbach, Matthieu Geist, Sergey Levine, Ruslan SalakhutdinovICML 2023 · 8 citations
- Understanding Policy Gradient Algorithms: A Sensitivity-Based ApproachShuang Wu, Ling Shi, Jun Wang, Guangjian TianICML 2022 · 7 citations
Builds on3
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 144 citations
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 137 citations
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 30 citations
Related papers
- Learning in Stackelberg Mean Field Games: A Non-Asymptotic AnalysisSihan Zeng, Benjamin Patrick Evans, Sujay Bhatt, Leo Ardon et al.NeurIPS 2025 · 1 citation
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 110 citations
- Decision-Aware Actor-Critic with Function Approximation and Theoretical GuaranteesSharan Vaswani, Amirreza Kazemi, Reza Babanezhad Harikandeh, Nicolas Le RouxNeurIPS 2023 · 6 citations
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy OptimizationPierluca D'Oro, Wojciech JaskowskiNeurIPS 2020 · 33 citations
- Learning Value Functions in Deep Policy Gradients using Residual VarianceYannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe PreuxICLR 2021 · 16 citations
