General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler Divergence
Lingwei Zhu, Zheng Chen, Matthew Schlegel, Martha White
Abstract
Many policy optimization approaches in reinforcement learning incorporate a Kullback-Leilbler (KL) divergence to the previous policy, to prevent the policy from changing too quickly. This idea was initially proposed in a seminal paper on Conservative Policy Iteration, with approximations given by algorithms like TRPO and Munchausen Value Iteration (MVI). We continue this line of work by investigating a generalized KL divergence-called the Tsallis KL divergence. Tsallis KL defined by the q-logarithm is a strict generalization, as q = 1 corresponds to the standard KL divergence; q > 1 provides a range of new options. We characterize the types of policies learned under the Tsallis KL, and motivate when q > 1 could be beneficial. To obtain a practical algorithm that incorporates Tsallis KL regularization, we extend MVI, which is one of the simplest approaches to incorporate KL regularization. We show that this generalized MVI(q) obtains significant improvements over the standard MVI(q = 1) across 35 Atari games. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 50 citations
- f-Divergence Variational InferenceNeng Wan, Dapeng Li, Naira HovakimyanNeurIPS 2020 · 13 citations
Related papers
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- Mutual Information Regularized Offline Reinforcement LearningXiao Ma, Bingyi Kang, Zhongwen Xu, Min Lin et al.NeurIPS 2023 · 14 citations
- On-Policy Deep Reinforcement Learning for the Average-Reward CriterionYiming Zhang, Keith W. RossICML 2021 · 59 citations
- Convex Regularization in Monte-Carlo Tree SearchTuan Dam, Carlo D'Eramo, Jan Peters, Joni PajarinenICML 2021 · 12 citations
- q-exponential family for policy optimizationLingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai et al.ICLR 2025
