Restoring balance: principled under/oversampling of data for optimal classification
Emanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi Monasson
Abstract
Class imbalance in real-world data poses a common bottleneck for machine learning tasks, since achieving good generalization on under-represented examples is often challenging. Mitigation strategies, such as under or oversampling the data depending on their abundances, are routinely proposed and tested empirically, but how they should adapt to the data statistics remains poorly understood. In this work, we determine exact analytical expressions of the generalization curves in the high-dimensional regime for linear classifiers (Support Vector Machines). We also provide a sharp prediction of the effects of under/oversampling strategies depending on class imbalance, first and second moments of the data, and the metrics of performance considered. We show that mixed strategies involving under and oversampling of data lead to performance improvement. Through numerical experiments, we show the relevance of our theoretical predictions on real datasets, on deeper architectures and with sampling strategies based on unsupervised probabilistic models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cf91694-c970-47fd-8de7-3d45b7789418Cited by top-tier papers3
- Improved Balanced Classification with Theoretically Grounded Loss FunctionsCorinna Cortes, Mehryar Mohri, Yutao ZhongNeurIPS 2025 · 19 citations
- A Theoretical Framework For Overfitting In Energy-based ModelingGiovanni Catania, Aurélien Decelle, Cyril Furtlehner, Beatriz SeoaneICML 2025
- Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced DataCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
Builds on13
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- Long- Tailed Recognition via Weight BalancingShaden Alshammari, Yu-Xiong Wang, Deva Ramanan, Shu KongCVPR 2022 · 133 citations
- The Role of Regularization in Classification of High-dimensional Noisy Gaussian MixtureFrancesca Mignacco, Florent Krzakala, Yue M. Lu, Pierfrancesco Urbani et al.ICML 2020 · 98 citations
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco et al.NeurIPS 2021 · 70 citations
Related papers
- A Statistical Theory of Overfitting for Imbalanced ClassificationJingyang Lyu, Kangjie Zhou, Yiqiao ZhongICLR 2026 · 4 citations
- On the Error Resistance of Hinge-Loss MinimizationKunal TalwarNeurIPS 2020 · 7 citations
- Why does Throwing Away Data Improve Worst-Group Error?Kamalika Chaudhuri, Kartik Ahuja, Martín Arjovsky, David Lopez-PazICML 2023 · 27 citations
- Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!Xu Yang, Shaowei Wang, Yi Li, Shaohua WangICSE 2023 · 19 citations
- Simplifying Neural Network Training Under Class ImbalanceRavid Shwartz-Ziv, Micah Goldblum, Yucen Lily Li, C. Bayan Bruss et al.NeurIPS 2023 · 45 citations
