On the Convergence of Smooth Regularized Approximate Value Iteration Schemes
Elena Smirnova, Elvis Dohmatob
摘要
Entropy regularization, smoothing of Q-values and neural network function approximator are key components of the state-of-the-art reinforcement learning (RL) algorithms, such as Soft Actor-Critic [1] . Despite the widespread use, the impact of these core techniques on the convergence of RL algorithms is not yet fully understood. In this work, we analyse these techniques from error propagation perspective using the approximate dynamic programming framework. In particular, our analysis shows that (1) value smoothing results in increased stability of the algorithm in exchange for slower convergence, (2) entropy regularization reduces overestimation errors at the cost of modifying the original problem, (3) we study a combination of these techniques that describes the Soft Actor-Critic algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Faster Deep Reinforcement Learning with Slower Online NetworkKavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim 等NeurIPS 2022 · 被引用 7 次
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement LearningJacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. KulkarniAAAI 2023 · 被引用 13 次
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 被引用 26 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
- Refined Analysis of Entropy-Regularized Actor-CriticSafwan Labbi, Paul Mangold, Daniil Tiapkin, Eric MoulinesICML 2026
- Finite-time Convergence Analysis of Actor-Critic with Evolving RewardRui Hu, Yu Chen, Longbo HuangICML 2026
