Leverage the Average: an Analysis of KL Regularization in Reinforcement Learning
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, Matthieu Geist
Abstract
Recent Reinforcement Learning (RL) algorithms making use of Kullback-Leibler (KL) regularization as a core component have shown outstanding performance. Yet, only little is understood theoretically about why KL regularization helps, so far. We study KL regularization within an approximate value iteration scheme and show that it implicitly averages q-values. Leveraging this insight, we provide a very strong performance bound, the very first to combine two desirable aspects: a linear dependency to the horizon (instead of quadratic) and an error propagation term involving an averaging e ect of the estimation errors (instead of an accumulation e ect). We also study the more general case of an additional entropy regularizer. The resulting abstract scheme encompasses many existing RL algorithms. Some of our assumptions do not hold with neural networks, so we complement this theoretical analysis with an extensive empirical study. ú Work done while at DeepMind. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Doubly Regularized Markov Decision Processes for Robust Reinforcement LearningYiting He, Zhishuai Liu, Pan XuICML 2026 · 213 citations
- Emergent Communication at ScaleRahma Chaabouni, Florian Strub, Florent Altché, Eugene Tarassov et al.ICLR 2022 · 65 citations
- Offline Reinforcement Learning as Anti-explorationShideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot et al.AAAI 2022 · 64 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk et al.NeurIPS 2022 · 27 citations
Builds on2
Related papers
- Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and PracticeToshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard et al.ICML 2023 · 4 citations
- Mirror Descent Actor Critic via Bounded Advantage LearningRyo IwakiICML 2026
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 27 citations
- General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler DivergenceLingwei Zhu, Zheng Chen, Matthew Schlegel, Martha WhiteNeurIPS 2023 · 4 citations
- On the Convergence of Smooth Regularized Approximate Value Iteration SchemesElena Smirnova, Elvis DohmatobNeurIPS 2020 · 8 citations
