Synthetic Tabular Data Generation for Imbalanced Classification: The Surprising Effectiveness of an Overlap Class
Annie D'souza, Swetha M, Sunita Sarawagi
Abstract
Handling imbalance in class distribution when building a classifier over tabular data has been a problem of long-standing interest. One popular approach is augmenting the training dataset with synthetically generated data. While classical augmentation techniques were limited to linear interpolation of existing minority class examples, recently higher capacity deep generative models are providing greater promise.
However, handling of imbalance in class distribution when building a deep generative model is also a challenging problem, that has not been studied as extensively as imbalanced classifier model training. We show that state-of-the-art deep generative models yield significantly lower-quality minority examples than majority examples. We propose a novel technique of converting the binary class labels to ternary class labels by introducing a class for the region where minority and majority distributions overlap. We show that just this pre-processing of the training set, significantly improves the quality of data generated spanning several state-of-the-art diffusion and GAN-based models. While training the classifier using synthetic data, we remove the overlap class from the training data and justify the reasons behind the enhanced accuracy. We perform extensive experiments on four real-life datasets, five different classifiers, and five generative models, demonstrating that our method enhances not only the synthesizer performance of state-of-the-art models but also the classifier performance
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9d72e16-e605-4125-8d77-50c0460fd955Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- M2m: Imbalanced Classification via Major-to-Minor TranslationJaehyung Kim, Jongheon Jeong, Jinwoo ShinCVPR 2020
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang et al.ICLR 2024 · 8 citations
- Generative Adversarial Minority OversamplingSankha Subhra Mullick, Shounak Datta, Swagatam DasICCV 2019 · 222 citations
- Data Augmentation with Diffusion for Open-Set Semi-Supervised LearningSeonghyun Ban, Heesan Kong, Kee-Eung KimNeurIPS 2024 · 4 citations
- Differential Privacy Under Class Imbalance: Methods and Empirical InsightsLucas Rosenblatt, Yuliia Lut, Ethan Turok, Marco Avella Medina et al.ICML 2025
