Leverage the Average: an Analysis of KL Regularization in Reinforcement Learning
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, Matthieu Geist
摘要
Recent Reinforcement Learning (RL) algorithms making use of Kullback-Leibler (KL) regularization as a core component have shown outstanding performance. Yet, only little is understood theoretically about why KL regularization helps, so far. We study KL regularization within an approximate value iteration scheme and show that it implicitly averages q-values. Leveraging this insight, we provide a very strong performance bound, the very first to combine two desirable aspects: a linear dependency to the horizon (instead of quadratic) and an error propagation term involving an averaging e ect of the estimation errors (instead of an accumulation e ect). We also study the more general case of an additional entropy regularizer. The resulting abstract scheme encompasses many existing RL algorithms. Some of our assumptions do not hold with neural networks, so we complement this theoretical analysis with an extensive empirical study. ú Work done while at DeepMind. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Doubly Regularized Markov Decision Processes for Robust Reinforcement LearningYiting He, Zhishuai Liu, Pan XuICML 2026 · 被引用 213 次
- Emergent Communication at ScaleRahma Chaabouni, Florian Strub, Florent Altché, Eugene Tarassov 等ICLR 2022 · 被引用 65 次
- Offline Reinforcement Learning as Anti-explorationShideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot 等AAAI 2022 · 被引用 64 次
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 被引用 47 次
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk 等NeurIPS 2022 · 被引用 27 次
它引用的顶会 Paper2
相关 Paper
- Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and PracticeToshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard 等ICML 2023 · 被引用 4 次
- Mirror Descent Actor Critic via Bounded Advantage LearningRyo IwakiICML 2026
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 被引用 27 次
- General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler DivergenceLingwei Zhu, Zheng Chen, Matthew Schlegel, Martha WhiteNeurIPS 2023 · 被引用 4 次
- On the Convergence of Smooth Regularized Approximate Value Iteration SchemesElena Smirnova, Elvis DohmatobNeurIPS 2020 · 被引用 8 次
