On the Convergence of Smooth Regularized Approximate Value Iteration Schemes
Elena Smirnova, Elvis Dohmatob
Abstract
Entropy regularization, smoothing of Q-values and neural network function approximator are key components of the state-of-the-art reinforcement learning (RL) algorithms, such as Soft Actor-Critic [1] . Despite the widespread use, the impact of these core techniques on the convergence of RL algorithms is not yet fully understood. In this work, we analyse these techniques from error propagation perspective using the approximate dynamic programming framework. In particular, our analysis shows that (1) value smoothing results in increased stability of the algorithm in exchange for slower convergence, (2) entropy regularization reduces overestimation errors at the cost of modifying the original problem, (3) we study a combination of these techniques that describes the Soft Actor-Critic algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 277f1be3-6393-4672-8371-9cb37ff79f4cCited by top-tier papers2
- Faster Deep Reinforcement Learning with Slower Online NetworkKavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim et al.NeurIPS 2022 · 7 citations
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 3 citations
Builds on1
Related papers
- Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement LearningJacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. KulkarniAAAI 2023 · 13 citations
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 26 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- Refined Analysis of Entropy-Regularized Actor-CriticSafwan Labbi, Paul Mangold, Daniil Tiapkin, Eric MoulinesICML 2026
- Finite-time Convergence Analysis of Actor-Critic with Evolving RewardRui Hu, Yu Chen, Longbo HuangICML 2026
