Improved Stochastic Optimization of LogSumExp
Egor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel Dvurechenskii
Abstract
The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new -divergence called the Safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Latent Space Robust Optimization of Neural Processes with Aligned Stratified Order-Statistic Loss ReductionQi Tao, Jiarong Wen, Jing Yang, Guanlin Wu et al.ICML 2026
- A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk MinimizationXiyuan Wei, Linli Zhou, Bokun Wang, Chih-Jen Lin et al.ICML 2026
Builds on16
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 281 citations
- Neural Optimal TransportAlexander Korotin, Daniil Selikhanovych, Evgeny BurnaevICLR 2023 · 151 citations
- Score-based Generative Neural Networks for Large-Scale Optimal TransportGrady Daniels, Tyler Maunu, Paul HandNeurIPS 2021 · 101 citations
- An Online Method for A Class of Distributionally Robust Optimization with Non-convex ObjectivesQi Qi, Zhishuai Guo, Yi Xu, Rong Jin et al.NeurIPS 2021 · 61 citations
Related papers
- Efficient Optimal Transport Algorithm by Accelerated Gradient DescentDongsheng An, Na Lei, Xiaoyin Xu, Xianfeng GuAAAI 2022 · 18 citations
- Distributionally Robust Optimization with Bias and Variance ReductionRonak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd HarchaouiICLR 2024 · 6 citations
- Regularized Optimal Transport is Ground Cost AdversarialFrançois-Pierre Paty, Marco CuturiICML 2020 · 33 citations
- Distributionally Robust Bayesian Optimization with φ-divergencesHisham Husain, Vu Nguyen, Anton van den HengelNeurIPS 2023 · 26 citations
- Bootstrap Your Uncertainty: Adaptive Robust Classification Driven by Optimal-TransportJiawei Huang, Minming Li, Hu DingNeurIPS 2025
