Revisiting Peng's Q(λ) for Modern Reinforcement Learning
Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rémi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel
Abstract
Off-policy multi-step reinforcement learning algorithms consist of conservative and non-conservative algorithms: the former actively cut traces, whereas the latter do not. Recently, Munos et al. (2016) proved the convergence of conservative algorithms to an optimal Q-function. In contrast, non-conservative algorithms are thought to be unsafe and have a limited or no theoretical guarantee. Nonetheless, recent studies have shown that non-conservative algorithms empirically outperform conservative ones. Motivated by the empirical results and the lack of theory, we carry out theoretical analyses of Peng's Q(), a representative example of non-conservative algorithms. We prove that it also converges to an optimal policy provided that the behavior policy slowly tracks a greedy policy in a way similar to conservative policy iteration. Such a result has been conjectured to be true but has not been proven. We also experiment with Peng's Q() in complex continuous control tasks, confirming that Peng's Q() often outperforms conservative algorithms despite its simplicity. These results indicate that Peng's Q(), which was thought to be unsafe, is a theoretically-sound and practically effective algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4889c31-c8da-4cc5-a9cd-2b841db25a90Cited by top-tier papers8
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 114 citations
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 38 citations
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 19 citations
- The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement LearningYunhao Tang, Rémi Munos, Mark Rowland, Bernardo Ávila Pires et al.NeurIPS 2022 · 16 citations
- Averaging n-step Returns Reduces Variance in Reinforcement LearningBrett Daley, Martha White, Marlos C. MachadoICML 2024 · 7 citations
Builds on2
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
Related papers
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 173 citations
- Iteratively Refined Behavior Regularization for Offline Reinforcement LearningYi Ma, Jianye Hao, Xiaohan Hu, Yan Zheng et al.NeurIPS 2024 · 11 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
