Evaluating model performance under worst-case subpopulations
Mike Li, Hongseok Namkoong, Shangzhou Xia
Abstract
The performance of ML models degrades when the training population is different from that seen under operation. Towards assessing distributional robustness, we study the worst-case performance of a model over all subpopulations of a given size, defined with respect to core attributes Z. This notion of robustness can consider arbitrary (continuous) attributes Z, and automatically accounts for complex intersectionality in disadvantaged groups. We develop a scalable yet principled two-stage estimation procedure that can evaluate the robustness of state-of-the-art models. We prove that our procedure enjoys several finite-sample convergence guarantees, including dimension-free convergence. Instead of overly conservative notions based on Rademacher complexities, our evaluation error depends on the dimension of Z only through the out-of-sample error in estimating the performance conditional on Z. On real datasets, we demonstrate that our method certifies the robustness of a model and prevents deployment of unreliable models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca048667-bc23-41f7-9c67-20267ebedf40Cited by top-tier papers6
- Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationQi Bi, Shaodi You, Theo GeversAAAI 2024 · 77 citations
- Evaluating Robustness to Dataset Shift via Parametric Robustness SetsNikolaj Thams, Michael Oberst, David A. SontagNeurIPS 2022 · 17 citations
- Being Right for Whose Right Reasons?Terne Sasha Thorn Jakobsen, Laura Cabello, Anders SøgaardACL 2023 · 8 citations
- Stability Evaluation through Distributional Perturbation AnalysisJosé H. Blanchet, Peng Cui, Jiajin Li, Jiashuo LiuICML 2024 · 6 citations
- Responsible AI (RAI) Games and EnsemblesYash Gupta, Runtian Zhai, Arun Suggala, Pradeep RavikumarNeurIPS 2023 · 1 citation
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
Related papers
- Multiply Robust Estimation for Local Distribution Shifts with Multiple DomainsSteven Wilkins-Reeves, Xu Chen, Qi Ma, Christine Agarwal et al.ICML 2024 · 2 citations
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
- MixMax: Distributional Robustness in Function Space via Optimal Data MixturesAnvith Thudi, Chris J. MaddisonICLR 2025
- Size-adaptive Hypothesis Testing for FairnessAntonio Ferrara, Francesco Cozzi, Alan Perotti, André Panisson et al.NeurIPS 2025 · 2 citations
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 51 citations
