Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-Authentication
June Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang, Eunho Yang
Abstract
Tabular data in the wild are frequently afflicted with class-imbalance, biasing machine learning model predictions towards major classes. A data-centric solution to this problem is oversampling -where the classes are balanced by adding synthetic minority samples via generative methods. However, although tabular generative models are capable of generating synthetic samples under a balanced distribution, their integrity suffers when the number of minority samples is low. To this end, pre-trained generative language models with rich prior knowledge are a fitting candidate for the task at hand. Nevertheless, an oversampling strategy tailored for tabular data that utilizes the extensive capabilities of such language models is yet to emerge. In this paper, we propose a novel oversampling framework for tabular data to channel the abilities of generative language models. By leveraging its conditional sampling capabilities, we synthesize minority samples by progressively masking the important features of the majority class samples and imputing them towards the minority distribution. To reduce the inclusion of imperfectly converted samples, we utilize the power of the language model itself to self-authenticate the labels of the samples generated by itself, sifting out ill-converted samples. Extensive experiments on a variety of datasets and imbalance ratios reveal that the proposed method successfully generates reliable minority samples to boost the performance of machine learning classifiers, even under heavy imbalance ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77787823-a1de-421b-b2ad-648d431ce1f3Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- SOS: Score-based Oversampling for Tabular DataJayoung Kim, Chaejeong Lee, Yehjin Shin, Sewon Park et al.KDD 2022 · 20 citations
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationSeulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun et al.CVPR 2022 · 199 citations
- Generative Table Pre-training Empowers Models for Tabular PredictionTianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian et al.EMNLP 2023 · 18 citations
- Synthetic Tabular Data Generation for Imbalanced Classification: The Surprising Effectiveness of an Overlap ClassAnnie D'souza, Swetha M, Sunita SarawagiAAAI 2025 · 9 citations
