Gradient Flow Sampler-based Distributionally Robust Optimization
Zusen Xu, Jia-Jie Zhu
Abstract
We propose a mathematically principled PDE gradient flow framework for distributionally robust optimization (DRO). Exploiting the recent advances in the intersection of Monte Carlo sampling and statistical optimal transport, we show that our theoretical framework can be implemented as practical algorithms for sampling from worst-case distributions and, consequently, DRO. While numerous previous works have relied on dual reformulation techniques, we contribute a sound and complete gradient flow view based on SDEs or PDEs that can be used to construct new algorithms for general, potentially non-convex, losses. Without loss of generality, we solve a class of Wasserstein and entropy-regularized DRO problems using the recently-discovered Wasserstein Fisher-Rao and Stein variational gradient flows. Notably, we also show some simple reductions of our framework recover exactly previously proposed popular DRO methods, and provide new insights into their theoretical limits and optimization dynamics of DRO. Numerical studies based on stochastic gradient descent on machine learning tasks provide empirical backing for our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd524f93-ef70-4f03-bcc1-a82f6ae1fc82Builds on6
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 281 citations
- Understanding the Variance Collapse of SVGD in High DimensionsJimmy Ba, Murat A. Erdogdu, Marzyeh Ghassemi, Shengyang Sun et al.ICLR 2022 · 35 citations
- A Finite-Particle Convergence Rate for Stein Variational Gradient DescentJiaxin Shi, Lester MackeyNeurIPS 2023 · 34 citations
- A Convergence Theory for SVGD in the Population Limit under Talagrand's Inequality T1Adil Salim, Lukang Sun, Peter RichtárikICML 2022 · 28 citations
- Strategic Distribution Shift of Interacting Agents via Coupled Gradient FlowsLauren E. Conger, Franca Hoffmann, Eric Mazumdar, Lillian J. RatliffNeurIPS 2023 · 7 citations
Related papers
- Bootstrap Your Uncertainty: Adaptive Robust Classification Driven by Optimal-TransportJiawei Huang, Minming Li, Hu DingNeurIPS 2025
- Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust OptimizationShuang Liu, Yihan Wang, Yifan Zhu, Yibo Miao et al.ICLR 2025
- Knowledge-Guided Wasserstein Distributionally Robust OptimizationZitao Wang, Ziyuan Wang, Molei Liu, Nian SiICML 2025
- Distributionally Robust Optimization with Data GeometryJiashuo Liu, Jiayun Wu, Bo Li, Peng CuiNeurIPS 2022 · 28 citations
- Non-convex Distributionally Robust Optimization: Non-asymptotic AnalysisJikai Jin, Bohang Zhang, Haiyang Wang, Liwei WangNeurIPS 2021 · 65 citations
