A Closer Look at AUROC and AUPRC under Class Imbalance
Matthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti, Jack Gallifant
Abstract
In machine learning (ML), a widespread claim is that the area under the precision-recall curve (AUPRC) is a superior metric for model comparison to the area under the receiver operating characteristic (AUROC) for tasks with class imbalance. This paper refutes this notion on two fronts. First, we theoretically characterize the behavior of AUROC and AUPRC in the presence of model mistakes, establishing clearly that AUPRC is not generally superior in cases of class imbalance. We further show that AUPRC can be a harmful metric as it can unduly favor model improvements in subpopulations with more frequent positive labels, heightening algorithmic disparities. Next, we empirically support our theory using experiments on both semi-synthetic and real-world fairness datasets. Prompted by these insights, we conduct a review of over 1.5 million scientific papers to understand the origin of this invalid claim, finding that it is often made without citation, misattributed to papers that do not argue this point, and aggressively over-generalized from source arguments. Our findings represent a dual contribution: a significant technical advancement in understanding the relationship between AUROC and AUPRC and a stark warning about unchecked assumptions in the ML community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent CollaborationZixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng et al.WWW 2025 · 34 citations
- Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow MatchingZhong Li, Qi Huang, Yuxuan Zhu, Lincen Yang et al.NeurIPS 2025 · 16 citations
- Benchmarking ECG FMs: A Reality Check Across Clinical TasksM A Al-Masud, Juan Lopez Alcaraz, Nils StrodthoffICLR 2026 · 9 citations
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu et al.ACM MM 2024 · 8 citations
- Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error CostsGerardo Flores, Alyssa H. Smith, Julia Fukuyama, Ashia C. WilsonNeurIPS 2025 · 5 citations
Builds on4
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Deep Weakly-supervised Anomaly DetectionGuansong Pang, Chunhua Shen, Huidong Jin, Anton van den HengelKDD 2023 · 100 citations
- Detecting Semantic AnomaliesFaruk Ahmed, Aaron C. CourvilleAAAI 2020 · 93 citations
- StatEcoNet: Statistical Ecology Neural Networks for Species Distribution ModelingEugene Seo, Rebecca A. Hutchinson, Xiao Fu, Chelsea Li et al.AAAI 2021 · 10 citations
Related papers
- Interplay of ROC and Precision-Recall AUCs: Theoretical Limits and Practical Implications in Binary ClassificationMartin Mihelich, François Castagnos, Charles DogninICML 2024
- Minimax AUC Fairness: Efficient Algorithm with Provable ConvergenceZhenhuan Yang, Yan Lok Ko, Kush R. Varshney, Yiming YingAAAI 2023 · 22 citations
- AUC Maximization under Positive Distribution ShiftAtsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama et al.NeurIPS 2024 · 7 citations
- A Unified Framework against Topology and Class ImbalanceJunyu Chen, Qianqian Xu, Zhiyong Yang, Xiaochun Cao et al.ACM MM 2022 · 4 citations
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 44 citations
