A Theoretical Analysis of the Learning Dynamics under Class Imbalance
Emanuele Francazi, Marco Baity-Jesi, Aurélien Lucchi
Abstract
Data imbalance is a common problem in machine learning that can have a critical effect on the performance of a model. Various solutions exist but their impact on the convergence of the learning dynamics is not understood. Here, we elucidate the significant negative impact of data imbalance on learning, showing that the learning curves for minority and majority classes follow sub-optimal trajectories when training with a gradient-based optimizer. This slowdown is related to the imbalance ratio and can be traced back to a competition between the optimization of different classes. Our main contribution is the analysis of the convergence of full-batch (GD) and stochastic gradient descent (SGD), and of variants that renormalize the contribution of each per-class gradient. We find that GD is not guaranteed to decrease the loss for each class but that this problem can be addressed by performing a per-class normalization of the gradient. With SGD, class imbalance has an additional effect on the direction of the gradients: the minority class suffers from a higher directional noise, which reduces the effectiveness of the per-class gradient normalization. Our findings not only allow us to understand the potential and limitations of strategies involving the per-class gradients, but also the reason for the effectiveness of previously used solutions for class imbalance such as oversampling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdfd2873-73dd-4e36-9391-1dd57d5a4486Cited by top-tier papers9
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt et al.NeurIPS 2024 · 100 citations
- Restoring balance: principled under/oversampling of data for optimal classificationEmanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi MonassonICML 2024 · 13 citations
- Initial Guessing Bias: How Untrained Networks Favor Some ClassesEmanuele Francazi, Aurélien Lucchi, Marco Baity-JesiICML 2024 · 7 citations
- When majority rules, minority loses: bias amplification of gradient descentFrançois Bachoc, Jérôme Bolte, Ryan Boustany, Jean-Michel LoubesNeurIPS 2025 · 3 citations
- Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated LearningHaemin Park, Diego Klabjan, Martin Braun, Xiuqi Li et al.ICML 2026
Builds on1
Related papers
- Procrustean Training for Imbalanced Deep LearningHan-Jia Ye, De-Chuan Zhan, Wei-Lun ChaoICCV 2021 · 36 citations
- Learn2Mix: Training Neural Networks Using Adaptive Data IntegrationShyam Venkatasubramanian, Vahid TarokhNeurIPS 2025 · 2 citations
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han et al.ICLR 2021 · 165 citations
- A Re-Balancing Strategy for Class-Imbalanced Classification Based on Instance DifficultySihao Yu, Jiafeng Guo, Ruqing Zhang, Yixing Fan et al.CVPR 2022 · 42 citations
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 1 citation
