Evaluating model performance under worst-case subpopulations
Mike Li, Hongseok Namkoong, Shangzhou Xia
摘要
The performance of ML models degrades when the training population is different from that seen under operation. Towards assessing distributional robustness, we study the worst-case performance of a model over all subpopulations of a given size, defined with respect to core attributes Z. This notion of robustness can consider arbitrary (continuous) attributes Z, and automatically accounts for complex intersectionality in disadvantaged groups. We develop a scalable yet principled two-stage estimation procedure that can evaluate the robustness of state-of-the-art models. We prove that our procedure enjoys several finite-sample convergence guarantees, including dimension-free convergence. Instead of overly conservative notions based on Rademacher complexities, our evaluation error depends on the dimension of Z only through the out-of-sample error in estimating the performance conditional on Z. On real datasets, we demonstrate that our method certifies the robustness of a model and prevents deployment of unreliable models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationQi Bi, Shaodi You, Theo GeversAAAI 2024 · 被引用 77 次
- Evaluating Robustness to Dataset Shift via Parametric Robustness SetsNikolaj Thams, Michael Oberst, David A. SontagNeurIPS 2022 · 被引用 17 次
- Being Right for Whose Right Reasons?Terne Sasha Thorn Jakobsen, Laura Cabello, Anders SøgaardACL 2023 · 被引用 8 次
- Stability Evaluation through Distributional Perturbation AnalysisJosé H. Blanchet, Peng Cui, Jiajin Li, Jiashuo LiuICML 2024 · 被引用 6 次
- Responsible AI (RAI) Games and EnsemblesYash Gupta, Runtian Zhai, Arun Suggala, Pradeep RavikumarNeurIPS 2023 · 被引用 1 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li 等CVPR 2022 · 被引用 364 次
相关 Paper
- Multiply Robust Estimation for Local Distribution Shifts with Multiple DomainsSteven Wilkins-Reeves, Xu Chen, Qi Ma, Christine Agarwal 等ICML 2024 · 被引用 2 次
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras 等NeurIPS 2025 · 被引用 9 次
- MixMax: Distributional Robustness in Function Space via Optimal Data MixturesAnvith Thudi, Chris J. MaddisonICLR 2025
- Size-adaptive Hypothesis Testing for FairnessAntonio Ferrara, Francesco Cozzi, Alan Perotti, André Panisson 等NeurIPS 2025 · 被引用 2 次
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 被引用 51 次
