A Theoretical Analysis of the Learning Dynamics under Class Imbalance
Emanuele Francazi, Marco Baity-Jesi, Aurélien Lucchi
摘要
Data imbalance is a common problem in machine learning that can have a critical effect on the performance of a model. Various solutions exist but their impact on the convergence of the learning dynamics is not understood. Here, we elucidate the significant negative impact of data imbalance on learning, showing that the learning curves for minority and majority classes follow sub-optimal trajectories when training with a gradient-based optimizer. This slowdown is related to the imbalance ratio and can be traced back to a competition between the optimization of different classes. Our main contribution is the analysis of the convergence of full-batch (GD) and stochastic gradient descent (SGD), and of variants that renormalize the contribution of each per-class gradient. We find that GD is not guaranteed to decrease the loss for each class but that this problem can be addressed by performing a per-class normalization of the gradient. With SGD, class imbalance has an additional effect on the direction of the gradients: the minority class suffers from a higher directional noise, which reduces the effectiveness of the per-class gradient normalization. Our findings not only allow us to understand the potential and limitations of strategies involving the per-class gradients, but also the reason for the effectiveness of previously used solutions for class imbalance such as oversampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt 等NeurIPS 2024 · 被引用 100 次
- Restoring balance: principled under/oversampling of data for optimal classificationEmanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi MonassonICML 2024 · 被引用 13 次
- Initial Guessing Bias: How Untrained Networks Favor Some ClassesEmanuele Francazi, Aurélien Lucchi, Marco Baity-JesiICML 2024 · 被引用 7 次
- When majority rules, minority loses: bias amplification of gradient descentFrançois Bachoc, Jérôme Bolte, Ryan Boustany, Jean-Michel LoubesNeurIPS 2025 · 被引用 3 次
- Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated LearningHaemin Park, Diego Klabjan, Martin Braun, Xiuqi Li 等ICML 2026
它引用的顶会 Paper1
相关 Paper
- Procrustean Training for Imbalanced Deep LearningHan-Jia Ye, De-Chuan Zhan, Wei-Lun ChaoICCV 2021 · 被引用 36 次
- Learn2Mix: Training Neural Networks Using Adaptive Data IntegrationShyam Venkatasubramanian, Vahid TarokhNeurIPS 2025 · 被引用 2 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
- A Re-Balancing Strategy for Class-Imbalanced Classification Based on Instance DifficultySihao Yu, Jiafeng Guo, Ruqing Zhang, Yixing Fan 等CVPR 2022 · 被引用 42 次
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 被引用 1 次
