Revisiting Peng's Q(λ) for Modern Reinforcement Learning
Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rémi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel
摘要
Off-policy multi-step reinforcement learning algorithms consist of conservative and non-conservative algorithms: the former actively cut traces, whereas the latter do not. Recently, Munos et al. (2016) proved the convergence of conservative algorithms to an optimal Q-function. In contrast, non-conservative algorithms are thought to be unsafe and have a limited or no theoretical guarantee. Nonetheless, recent studies have shown that non-conservative algorithms empirically outperform conservative ones. Motivated by the empirical results and the lack of theory, we carry out theoretical analyses of Peng's Q(), a representative example of non-conservative algorithms. We prove that it also converges to an optimal policy provided that the behavior policy slowly tracks a greedy policy in a way similar to conservative policy iteration. Such a result has been conjectured to be true but has not been proven. We also experiment with Peng's Q() in complex continuous control tasks, confirming that Peng's Q() often outperforms conservative algorithms despite its simplicity. These results indicate that Peng's Q(), which was thought to be unsafe, is a theoretically-sound and practically effective algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 被引用 38 次
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement LearningYunhao Tang, Rémi Munos, Mark Rowland, Bernardo Ávila Pires 等NeurIPS 2022 · 被引用 16 次
- Averaging n-step Returns Reduces Variance in Reinforcement LearningBrett Daley, Martha White, Marlos C. MachadoICML 2024 · 被引用 7 次
它引用的顶会 Paper2
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
相关 Paper
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 被引用 173 次
- Iteratively Refined Behavior Regularization for Offline Reinforcement LearningYi Ma, Jianye Hao, Xiaohan Hu, Yan Zheng 等NeurIPS 2024 · 被引用 11 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
