Efficient Augmentation for Imbalanced Deep Learning
Damien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. Chawla
摘要
Deep learning models may not effectively generalize across under-represented or minority classes. We empirically study a convolutional neural network’s (CNN) internal representation of imbalanced image data and measure the generalization gap between a model’s feature embeddings in the training and test sets, showing that the gap is wider for minority classes. This insight enables us to design an efficient three-phase CNN training framework for imbalanced data. The framework involves training the network end-to-end on imbalanced data to learn feature embeddings, performing data augmentation in the learned embedding space to balance the training data distribution, and fine-tuning the classifier head on the embedded balanced training data. We develop Expansive Over-Sampling (EOS) as a data augmentation technique to utilize in the training framework. EOS forms synthetic training instances as convex combinations between the minority class samples and their nearest adversaries in the embedding space to reduce the generalization gap. The proposed framework improves the accuracy over leading cost-sensitive and resampling methods commonly used in imbalanced learning. Moreover, it is more computationally efficient than standard data pre-processing methods, such as SMOTE and GAN-based over-sampling, as it requires fewer parameters and less training time. The source code for the proposed framework is available at: https://github.com/dd1github/EOS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy 等ICCV 2021 · 被引用 778 次
- Generative Adversarial Minority OversamplingSankha Subhra Mullick, Shounak Datta, Swagatam DasICCV 2019 · 被引用 222 次
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 被引用 28 次
- Stochastic smoothing of the top-K calibrated hinge loss for deep imbalanced classificationCamille Garcin, Maximilien Servajean, Alexis Joly, Joseph SalmonICML 2022 · 被引用 14 次
相关 Paper
- M2m: Imbalanced Classification via Major-to-Minor TranslationJaehyung Kim, Jongheon Jeong, Jinwoo ShinCVPR 2020
- Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text ClassificationJiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng 等EMNLP 2021 · 被引用 13 次
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng 等ICML 2022 · 被引用 46 次
- Domain Gap Embeddings for Generative Dataset AugmentationYinong Oliver Wang, Younjoon Chung, Chen Henry Wu, Fernando De la TorreCVPR 2024 · 被引用 8 次
- The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationSeulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun 等CVPR 2022 · 被引用 199 次
