Selective Classification Can Magnify Disparities Across Groups
Erik Jones, Shiori Sagawa, Pang Wei Koh, Ananya Kumar, Percy Liang
Abstract
Selective classification, in which models are allowed to abstain on uncertain predictions, is a natural approach to improving accuracy in settings where errors are costly but abstentions are manageable. In this paper, we find that while selective classification can improve average accuracies, it can simultaneously magnify existing accuracy disparities between various groups within a population, especially in the presence of spurious correlations. We observe this behavior consistently across five datasets from computer vision and NLP. Surprisingly, increasing the abstention rate can even decrease accuracies on some groups. To better understand when selective classification improves or worsens accuracy on a group, we study its margin distribution, which captures the model's confidences over all predictions. For example, when the margin distribution is symmetric, we prove that whether selective classification monotonically improves or worsens accuracy is fully determined by the accuracy at full coverage (i.e., without any abstentions) and whether the distribution satisfies a property we term left-log-concavity. Our analysis also shows that selective classification tends to magnify accuracy disparities that are present at full coverage. Fortunately, we find that it uniformly improves each group when applied to distributionally-robust models that achieve similar full-coverage accuracies across groups. Altogether, our results imply selective classification should be used with care and underscore the importance of models that perform equally well across groups at full coverage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b189f8a-4251-4416-baca-ff8be461dfbbCited by top-tier papers18
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Selective Regression under Fairness CriteriaAbhin Shah, Yuheng Bu, Joshua K. Lee, Subhro Das et al.ICML 2022 · 39 citations
- Fair Selective Classification Via SufficiencyJoshua K. Lee, Yuheng Bu, Deepta Rajan, Prasanna Sattigeri et al.ICML 2021 · 33 citations
- Selective Ensembles for Consistent PredictionsEmily Black, Klas Leino, Matt FredriksonICLR 2022 · 29 citations
- Learning to Defer with Limited Expert PredictionsPatrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik et al.AAAI 2023 · 28 citations
Builds on4
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Regression under Human AssistanceAbir De, Paramita Koley, Niloy Ganguly, Manuel Gomez-RodriguezAAAI 2020 · 73 citations
Related papers
- Selective Omniprediction and Fair AbstentionSílvia Casacuberta, Varun KanadeNeurIPS 2025 · 3 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 6 citations
- Distributionally Robust Optimization with Probabilistic GroupSoumya Suvra Ghosal, Yixuan LiAAAI 2023 · 14 citations
- Label-Efficient Group Robustness via Out-of-Distribution Concept CurationYiwei Yang, Anthony Z. Liu, Robert Wolfe, Aylin Caliskan et al.CVPR 2024
