Generalization Bounds with Minimal Dependency on Hypothesis Class via Distributionally Robust Optimization
Yibo Zeng, Henry Lam
Abstract
Established approaches to obtain generalization bounds in data-driven optimization and machine learning mostly build on solutions from empirical risk minimization (ERM), which depend crucially on the functional complexity of the hypothesis class. In this paper, we present an alternate route to obtain these bounds on the solution from distributionally robust optimization (DRO), a recent data-driven optimization framework based on worst-case analysis and the notion of ambiguity set to capture statistical uncertainty. In contrast to the hypothesis class complexity in ERM, our DRO bounds depend on the ambiguity set geometry and its compatibility with the true loss function. Notably, when using statistical distances such as maximum mean discrepancy, Wasserstein distance, or -divergence in the DRO, our analysis implies generalization bounds whose dependence on the hypothesis class appears the minimal possible: The bound depends solely on the true loss function, independent of any other candidates in the hypothesis class. To our best knowledge, it is the first generalization bound of this type in the literature, and we hope our findings can open the door for a better understanding of DRO, especially its benefits on loss minimization and other machine learning applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Decision-Focused Learning with Directional GradientsMichael Huang, Vishal GuptaNeurIPS 2024 · 25 citations
- Exact Generalization Guarantees for (Regularized) Wasserstein Distributionally Robust ModelsWaïss Azizian, Franck Iutzeler, Jérôme MalickNeurIPS 2023 · 14 citations
- Distributionally Robust Optimization with Bias and Variance ReductionRonak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd HarchaouiICLR 2024 · 6 citations
- Universal generalization guarantees for Wasserstein distributionally robust modelsTam Le, Jérôme MalickICLR 2025
- Certifiably Robust Model Evaluation in Federated Learning under Meta-Distributional ShiftsAmir Najafi, Samin Mahdizadeh Sani, Farzan FarniaICML 2025
Builds on3
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 134 citations
- Stable Adversarial Learning under Distributional ShiftsJiashuo Liu, Zheyan Shen, Peng Cui, Linjun Zhou et al.AAAI 2021 · 36 citations
Related papers
- Distributionally Robust Optimization via Generative Ambiguity ModelingJiaqi Wen, Jianyi YangICLR 2026 · 3 citations
- Outlier-Robust Wasserstein DROSloan Nietert, Ziv Goldfeld, Soroosh ShafieeNeurIPS 2023 · 26 citations
- Distributionally Robust Optimization with Data GeometryJiashuo Liu, Jiayun Wu, Bo Li, Peng CuiNeurIPS 2022 · 28 citations
- Outlier-Robust Distributionally Robust Optimization via Unbalanced Optimal TransportZifan Wang, Yi Shen, Michael M. Zavlanos, Karl Henrik JohanssonNeurIPS 2024 · 16 citations
- Geometry-Calibrated DRO: Combating Over-Pessimism with Free Energy ImplicationsJiashuo Liu, Jiayun Wu, Tianyu Wang, Hao Zou et al.ICML 2024 · 5 citations
