MixMax: Distributional Robustness in Function Space via Optimal Data Mixtures
Anvith Thudi, Chris J. Maddison
Abstract
Machine learning models are often required to perform well across several predefined settings, such as a set of user groups. Worst-case performance is a common metric to capture this requirement, and is the objective of group distributionally robust optimization (group DRO). Unfortunately, these methods struggle when the loss is non-convex in the parameters, or the model class is non-parametric. Here, we make a classical move to address this: we reparameterize group DRO from parameter space to function space, which results in a number of advantages. First, we show that group DRO over the space of bounded functions admits a minimax theorem. Second, for cross-entropy and mean squared error, we show that the minimax optimal mixture distribution is the solution of a simple convex optimization problem. Thus, provided one is working with a model class of universal function approximators, group DRO can be solved by a convex optimization problem followed by a classical risk minimization problem. We call our method MixMax. In our experiments, we found that MixMax matched or outperformed the standard group DRO baselines, and in particular, MixMax improved the performance of XGBoost over the only baseline, data balancing, for variations of the ACSIncome and CelebA annotations datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf5ec3c6-53bc-46ab-a002-519f23840f7dBuilds on4
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
- Towards a statistical theory of data selection under weak supervisionGermain Kolossov, Andrea Montanari, Pulkit TandonICLR 2024 · 27 citations
Related papers
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Geometry-Calibrated DRO: Combating Over-Pessimism with Free Energy ImplicationsJiashuo Liu, Jiayun Wu, Tianyu Wang, Hao Zou et al.ICML 2024 · 5 citations
- Multi-Expert Distributionally Robust Optimization for Out-of-Distribution GeneralizationJinyong Jeong, Hyungu Kahng, Seoung Bum KimNeurIPS 2025 · 6 citations
- Distributionally Robust Optimization with Probabilistic GroupSoumya Suvra Ghosal, Yixuan LiAAAI 2023 · 14 citations
- Non-convex Distributionally Robust Optimization: Non-asymptotic AnalysisJikai Jin, Bohang Zhang, Haiyang Wang, Liwei WangNeurIPS 2021 · 65 citations
