Modeling the Second Player in Distributionally Robust Optimization
Paul Michel, Tatsunori Hashimoto, Graham Neubig
Abstract
Distributionally robust optimization (DRO) provides a framework for training machine learning models that are able to perform well on a collection of related data distributions (the "uncertainty set"). This is done by solving a min-max game: the model is trained to minimize its maximum expected loss among all distributions in the uncertainty set. While careful design of the uncertainty set is critical to the success of the DRO procedure, previous work has been limited to relatively simple alternatives that keep the min-max optimization problem exactly tractable, such as f -divergence balls. In this paper, we argue instead for the use of neural generative models to characterize the worst-case distribution, allowing for more flexible and problem-specific selection of the uncertainty set. However, while simple conceptually, this approach poses a number of implementation and optimization challenges. To circumvent these issues, we propose a relaxation of the KL-constrained inner maximization objective that makes the DRO problem more amenable to gradient-based optimization of large scale generative models, and develop model selection heuristics to guide hyper-parameter search. On both toy settings and realistic NLP tasks, we find that the proposed approach yields models that are more robust than comparable baselines 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6c38466-71a3-428f-9dfa-db7e18217b00Cited by top-tier papers21
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 78 citations
- DORO: Distributional and Outlier Robust OptimizationRuntian Zhai, Chen Dan, J. Zico Kolter, Pradeep RavikumarICML 2021 · 74 citations
- Learning to Augment Distributions for Out-of-distribution DetectionQizhou Wang, Zhen Fang, Yonggang Zhang, Feng Liu et al.NeurIPS 2023 · 59 citations
- Understanding Contrastive Learning via Distributionally Robust OptimizationJunkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi et al.NeurIPS 2023 · 55 citations
- UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupZongbo Han, Zhipeng Liang, Fan Yang, Liu Liu et al.NeurIPS 2022 · 53 citations
Builds on5
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova et al.AAAI 2020 · 48 citations
- Distributional Robustness with IPMs and links to Regularization and GANsHisham HusainNeurIPS 2020 · 25 citations
- Robust Bayesian Classification Using An Optimistic Score RatioViet Anh Nguyen, Nian Si, Jose H. BlanchetICML 2020 · 15 citations
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 3 citations
Related papers
- Distributionally Robust Optimization via Generative Ambiguity ModelingJiaqi Wen, Jianyi YangICLR 2026 · 3 citations
- Distributionally Robust Models with Parametric Likelihood RatiosPaul Michel, Tatsunori Hashimoto, Graham NeubigICLR 2022 · 21 citations
- An Online Method for A Class of Distributionally Robust Optimization with Non-convex ObjectivesQi Qi, Zhishuai Guo, Yi Xu, Rong Jin et al.NeurIPS 2021 · 61 citations
- Generalization Bounds with Minimal Dependency on Hypothesis Class via Distributionally Robust OptimizationYibo Zeng, Henry LamNeurIPS 2022 · 11 citations
- Outlier-Robust Distributionally Robust Optimization via Unbalanced Optimal TransportZifan Wang, Yi Shen, Michael M. Zavlanos, Karl Henrik JohanssonNeurIPS 2024 · 16 citations
