A Closer Look at AUROC and AUPRC under Class Imbalance
Matthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti, Jack Gallifant
摘要
In machine learning (ML), a widespread claim is that the area under the precision-recall curve (AUPRC) is a superior metric for model comparison to the area under the receiver operating characteristic (AUROC) for tasks with class imbalance. This paper refutes this notion on two fronts. First, we theoretically characterize the behavior of AUROC and AUPRC in the presence of model mistakes, establishing clearly that AUPRC is not generally superior in cases of class imbalance. We further show that AUPRC can be a harmful metric as it can unduly favor model improvements in subpopulations with more frequent positive labels, heightening algorithmic disparities. Next, we empirically support our theory using experiments on both semi-synthetic and real-world fairness datasets. Prompted by these insights, we conduct a review of over 1.5 million scientific papers to understand the origin of this invalid claim, finding that it is often made without citation, misattributed to papers that do not argue this point, and aggressively over-generalized from source arguments. Our findings represent a dual contribution: a significant technical advancement in understanding the relationship between AUROC and AUPRC and a stark warning about unchecked assumptions in the ML community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent CollaborationZixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng 等WWW 2025 · 被引用 34 次
- Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow MatchingZhong Li, Qi Huang, Yuxuan Zhu, Lincen Yang 等NeurIPS 2025 · 被引用 16 次
- Benchmarking ECG FMs: A Reality Check Across Clinical TasksM A Al-Masud, Juan Lopez Alcaraz, Nils StrodthoffICLR 2026 · 被引用 9 次
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu 等ACM MM 2024 · 被引用 8 次
- Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error CostsGerardo Flores, Alyssa H. Smith, Julia Fukuyama, Ashia C. WilsonNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper4
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee 等NeurIPS 2020 · 被引用 406 次
- Deep Weakly-supervised Anomaly DetectionGuansong Pang, Chunhua Shen, Huidong Jin, Anton van den HengelKDD 2023 · 被引用 100 次
- Detecting Semantic AnomaliesFaruk Ahmed, Aaron C. CourvilleAAAI 2020 · 被引用 93 次
- StatEcoNet: Statistical Ecology Neural Networks for Species Distribution ModelingEugene Seo, Rebecca A. Hutchinson, Xiao Fu, Chelsea Li 等AAAI 2021 · 被引用 10 次
相关 Paper
- Interplay of ROC and Precision-Recall AUCs: Theoretical Limits and Practical Implications in Binary ClassificationMartin Mihelich, François Castagnos, Charles DogninICML 2024
- Minimax AUC Fairness: Efficient Algorithm with Provable ConvergenceZhenhuan Yang, Yan Lok Ko, Kush R. Varshney, Yiming YingAAAI 2023 · 被引用 22 次
- AUC Maximization under Positive Distribution ShiftAtsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama 等NeurIPS 2024 · 被引用 7 次
- A Unified Framework against Topology and Class ImbalanceJunyu Chen, Qianqian Xu, Zhiyong Yang, Xiaochun Cao 等ACM MM 2022 · 被引用 4 次
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 被引用 44 次
