Well-Posed KL-Regularized Control via Wasserstein and Kalman–Wasserstein KL Divergences
Viktor Stein, Adwait Datar, Nihat Ay
Abstract
Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher–Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. On a double integrator and a cart-pole example, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db8a9ceb-1e02-4154-b881-ec04c4bfe720Builds on3
- KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint SupportPierre Glaser, Michael Arbel, Arthur GrettonNeurIPS 2021 · 49 citations
- On Pathologies in KL-Regularized Reinforcement Learning from Expert DemonstrationsTim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal et al.NeurIPS 2021 · 33 citations
- Accurate Quantization of Measures via Interacting Particle-based OptimizationLantian Xu, Anna Korba, Dejan SlepcevICML 2022 · 18 citations
Related papers
- Statistical and Geometrical properties of the Kernel Kullback-Leibler divergenceAnna Korba, Francis R. Bach, Clémentine ChazalNeurIPS 2024 · 5 citations
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo et al.ICLR 2026 · 2 citations
- Improved Stochastic Optimization of LogSumExpEgor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel DvurechenskiiICML 2026 · 3 citations
- Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement LearningChen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao et al.ICML 2025
- What Is It Like to Be a Noise? An Entropy-based Gaussian Noise Regularization for Diffusion ModelsPascal Chang, Kai Lascheit, Jingwei Tang, Markus Gross et al.CVPR 2026
