Smoothing Advantage Learning
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
摘要
Advantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the method tends to be unstable in the case of function approximation. In this paper, we propose a simple variant of AL, named smoothing advantage learning (SAL), to alleviate this problem. The key to our method is to replace the original Bellman Optimal operator in AL with a smooth one so as to obtain more reliable estimation of the temporal difference target. We give a detailed account of the resulting action gap and the performance bound for approximate SAL. Further theoretical analysis reveals that the proposed value smoothing technique not only helps to stabilize the training procedure of AL by controlling the trade-off between convergence rate and the upper bound of the approximation errors, but is beneficial to increase the action gap between the optimal and sub-optimal action value as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Stabilizing Q Learning Via Soft Mellowmax OperatorYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2021 · 被引用 10 次
- On the Convergence of Smooth Regularized Approximate Value Iteration SchemesElena Smirnova, Elvis DohmatobNeurIPS 2020 · 被引用 8 次
相关 Paper
- Robust Action Gap Increasing with Clipped Advantage LearningZhe Zhang, Yaozhong Gan, Xiaoyang TanAAAI 2022 · 被引用 3 次
- Direct Advantage EstimationHsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard SchölkopfNeurIPS 2022 · 被引用 20 次
- Mirror Descent Actor Critic via Bounded Advantage LearningRyo IwakiICML 2026
- An Implicit Trust Region Approach to Behavior Regularized Offline Reinforcement LearningZhe Zhang, Xiaoyang TanAAAI 2024 · 被引用 10 次
- Action Gaps and Advantages in Continuous-Time Distributional Reinforcement LearningHarley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto 等NeurIPS 2024 · 被引用 9 次
