Efficient Augmentation for Imbalanced Deep Learning
Damien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. Chawla
Abstract
Deep learning models may not effectively generalize across under-represented or minority classes. We empirically study a convolutional neural network’s (CNN) internal representation of imbalanced image data and measure the generalization gap between a model’s feature embeddings in the training and test sets, showing that the gap is wider for minority classes. This insight enables us to design an efficient three-phase CNN training framework for imbalanced data. The framework involves training the network end-to-end on imbalanced data to learn feature embeddings, performing data augmentation in the learned embedding space to balance the training data distribution, and fine-tuning the classifier head on the embedded balanced training data. We develop Expansive Over-Sampling (EOS) as a data augmentation technique to utilize in the training framework. EOS forms synthetic training instances as convex combinations between the minority class samples and their nearest adversaries in the embedding space to reduce the generalization gap. The proposed framework improves the accuracy over leading cost-sensitive and resampling methods commonly used in imbalanced learning. Moreover, it is more computationally efficient than standard data pre-processing methods, such as SMOTE and GAN-based over-sampling, as it requires fewer parameters and less training time. The source code for the proposed framework is available at: https://github.com/dd1github/EOS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Generative Adversarial Minority OversamplingSankha Subhra Mullick, Shounak Datta, Swagatam DasICCV 2019 · 222 citations
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 28 citations
- Stochastic smoothing of the top-K calibrated hinge loss for deep imbalanced classificationCamille Garcin, Maximilien Servajean, Alexis Joly, Joseph SalmonICML 2022 · 14 citations
Related papers
- M2m: Imbalanced Classification via Major-to-Minor TranslationJaehyung Kim, Jongheon Jeong, Jinwoo ShinCVPR 2020
- Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text ClassificationJiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng et al.EMNLP 2021 · 13 citations
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng et al.ICML 2022 · 46 citations
- Domain Gap Embeddings for Generative Dataset AugmentationYinong Oliver Wang, Younjoon Chung, Chen Henry Wu, Fernando De la TorreCVPR 2024 · 8 citations
- The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationSeulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun et al.CVPR 2022 · 199 citations
