Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization
Shuang Liu, Yihan Wang, Yifan Zhu, Yibo Miao, Xiao-Shan Gao
Abstract
Wasserstein distributionally robust optimization (WDRO) optimizes against worst-case distributional shifts within a specified uncertainty set, leading to enhanced generalization on unseen adversarial examples, compared to standard adversarial training which focuses on pointwise adversarial perturbations. However, WDRO still suffers fundamentally from the robust overfitting problem, as it does not consider statistical error. We address this gap by proposing a novel robust optimization framework under a new uncertainty set for adversarial noise via Wasserstein distance and statistical error via Kullback-Leibler divergence, called the Statistically Robust WDRO. We establish a robust generalization bound for the new optimization framework, implying that out-of-distribution adversarial performance is at least as good as the statistically robust training loss with high probability. Furthermore, we derive conditions under which Stackelberg and Nash equilibria exist between the learner and the adversary, giving an optimal robust model in certain sense.Finally, through extensive experiments, we demonstrate that our method significantly mitigates robust overfitting and enhances robustness within the framework of WDRO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3057c59b-40f7-4e94-8b61-9bcb4bb8861eCited by top-tier papers2
- WILD-Diffusion: A WDRO Inspired Training Method for Diffusion Models under Limited DataXianglu Wang, Wanlin Zhang, Hu DingICLR 2026
- Distributionally Robust Set Representation Learning Under Inference-Time Element CorruptionYankai Chen, Hanrong Zhang, Bowei He, Philip Yu et al.ICML 2026
Builds on16
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
Related papers
- Outlier-Robust Wasserstein DROSloan Nietert, Ziv Goldfeld, Soroosh ShafieeNeurIPS 2023 · 26 citations
- Geometry-Calibrated DRO: Combating Over-Pessimism with Free Energy ImplicationsJiashuo Liu, Jiayun Wu, Tianyu Wang, Hao Zou et al.ICML 2024 · 5 citations
- Knowledge-Guided Wasserstein Distributionally Robust OptimizationZitao Wang, Ziyuan Wang, Molei Liu, Nian SiICML 2025
- Wasserstein Distributional Normalization For Robust Distributional Certification of Noisy Labeled DataSung Woo Park, Junseok KwonICML 2021 · 4 citations
- Gradient Flow Sampler-based Distributionally Robust OptimizationZusen Xu, Jia-Jie ZhuICML 2026
