Well-Posed KL-Regularized Control via Wasserstein and Kalman–Wasserstein KL Divergences
Viktor Stein, Adwait Datar, Nihat Ay
摘要
Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher–Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. On a double integrator and a cart-pole example, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint SupportPierre Glaser, Michael Arbel, Arthur GrettonNeurIPS 2021 · 被引用 49 次
- On Pathologies in KL-Regularized Reinforcement Learning from Expert DemonstrationsTim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal 等NeurIPS 2021 · 被引用 33 次
- Accurate Quantization of Measures via Interacting Particle-based OptimizationLantian Xu, Anna Korba, Dejan SlepcevICML 2022 · 被引用 18 次
相关 Paper
- Statistical and Geometrical properties of the Kernel Kullback-Leibler divergenceAnna Korba, Francis R. Bach, Clémentine ChazalNeurIPS 2024 · 被引用 5 次
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo 等ICLR 2026 · 被引用 2 次
- Improved Stochastic Optimization of LogSumExpEgor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel DvurechenskiiICML 2026 · 被引用 3 次
- Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement LearningChen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao 等ICML 2025
- What Is It Like to Be a Noise? An Entropy-based Gaussian Noise Regularization for Diffusion ModelsPascal Chang, Kai Lascheit, Jingwei Tang, Markus Gross 等CVPR 2026
