Value Improved Actor Critic Algorithms
Yaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok, Wendelin Boehmer, Matthijs T. J. Spaan
Abstract
To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greedification operators to iteratively improve it. The reliance on DNNs suggests an improvement that is gradient based, which is per step much less greedy than the improvement possible by greedier operators such as the greedy update used by Q-learning algorithms. On the other hand, slow changes to the policy can also be beneficial for the stability of the learning process, resulting in a tradeoff between greedification and stability. To better address this tradeoff, we propose to decouple the acting policy from the policy evaluated by the critic. This allows the agent to separately improve the critic's policy (e.g. value improvement) with greedier updates while maintaining the slow gradient-based improvement to the parameterized acting policy. We investigate the convergence of this approach in the finite-horizon domain using a popular analysis scheme which generalizes Policy Iteration with arbitrary improvement operators and approximate evaluation. Empirically, incorporating value-improvement into the popular off-policy actorcritic algorithms TD3 and SAC significantly improves or matches performance over the baselines respectively, across different environments from the DeepMind continuous control domain, with negligible compute and implementation cost 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 138 citations
- For SALE: State-Action Representation Learning for Deep Reinforcement LearningScott Fujimoto, Wei-Di Chang, Edward J. Smith, Shixiang Gu et al.NeurIPS 2023 · 128 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
Related papers
- Online Meta-Critic Learning for Off-Policy Actor-Critic MethodsWei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang et al.NeurIPS 2020 · 54 citations
- Decision-Aware Actor-Critic with Function Approximation and Theoretical GuaranteesSharan Vaswani, Amirreza Kazemi, Reza Babanezhad Harikandeh, Nicolas Le RouxNeurIPS 2023 · 6 citations
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 116 citations
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement LearningHiroki Furuta, Tadashi Kozuno, Tatsuya Matsushima, Yutaka Matsuo et al.NeurIPS 2021 · 14 citations
- Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningMotoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta et al.ICML 2025
