A Statistical Theory of Overfitting for Imbalanced Classification
Jingyang Lyu, Kangjie Zhou, Yiqiao Zhong
Abstract
Classification with imbalanced data is a common challenge in machine learning, where minority classes form only a small fraction of the training samples. Classical theory, relying on large-sample asymptotics and finite-sample corrections, is often ineffective in high dimensions, leaving many overfitting phenomena unexplained. In this paper, we develop a statistical theory for high-dimensional imbalanced linear classification, showing that dimensionality induces truncation or skewing effects on the logit distribution, which we characterize via a variational problem. For linearly separable Gaussian mixtures, logits follow on the test set but converge to on the training set---a pervasive phenomenon we confirm on tabular, image, and text data. This phenomenon explains why the minority class is more severely affected by overfitting. We further show that margin rebalancing mitigates minority accuracy drop and provide theoretical insights into calibration and uncertainty quantification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Thumb on the Scale: Optimal Loss Weighting in Last Layer RetrainingNathan Stromberg, Christos Thrampoulidis, Lalitha SankarNeurIPS 2025 · 1 citation
- Softmax is not Enough (for Adaptive Conformal Classification)Navid Akhavan Attar, Hesam Asadollahzadeh, Ling Luo, Uwe AickelinICLR 2026
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- Label-Imbalanced and Group-Sensitive Classification under OverparameterizationGanesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, Christos ThrampoulidisNeurIPS 2021 · 122 citations
- Distribution-free binary classification: prediction sets, confidence intervals and calibrationChirag Gupta, Aleksandr Podkopaev, Aaditya RamdasNeurIPS 2020 · 105 citations
Related papers
- Restoring balance: principled under/oversampling of data for optimal classificationEmanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi MonassonICML 2024 · 13 citations
- Why does Throwing Away Data Improve Worst-Group Error?Kamalika Chaudhuri, Kartik Ahuja, Martín Arjovsky, David Lopez-PazICML 2023 · 27 citations
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 57 citations
- QuanDA: Quantile-Based Discriminant Analysis for High-Dimensional Imbalanced ClassificationQian Tang, Yuwen Gu, Boxiang WangNeurIPS 2025
- Long-tailed Visual Recognition via Gaussian Clouded Logit AdjustmentMengke Li, Yiu-Ming Cheung, Yang LuCVPR 2022 · 84 citations
