EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
Jinhee Kim, Taesung Kim, Jaegul Choo
Abstract
Large language models (LLMs) have demonstrated remarkable in-context learning capabilities across diverse applications. In this work, we explore the effectiveness of LLMs for generating realistic synthetic tabular data, identifying key prompt design elements to optimize performance. We introduce EPIC, a novel approach that leverages balanced, grouped data samples and consistent formatting with unique variable mapping to guide LLMs in generating accurate synthetic data across all classes, even for imbalanced datasets. Evaluations on real-world datasets show that EPIC achieves state-of-the-art machine learning classification performance, significantly improving generation efficiency. These findings highlight the effectiveness of EPIC for synthetic tabular data generation, particularly in addressing class imbalance. Our source code for our work is available at: https://seharanul17.github.io/project-synthetic-tabular-llm/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c54a58f-6cde-4ca2-adc4-9db3e608f84eCited by top-tier papers2
- SAGE: Sparse Adaptive Guidance for Dependency-Aware Tabular Data GenerationShuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji KasneciACL 2026
- AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text ColumnsYuhao Zhang, Liang Yan, Shaoming Duan, Xinyu Zha et al.ACL 2026
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
Related papers
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang et al.ICLR 2024 · 8 citations
- Generative Table Pre-training Empowers Models for Tabular PredictionTianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian et al.EMNLP 2023 · 18 citations
- PANGEA: Projection-Based Augmentation with Non-Relevant General Data for Enhanced Domain Adaptation in LLMsSeungyoo Lee, Giung Nam, Moonseok Choi, Hyungi Lee et al.NeurIPS 2025
- SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisBangbang Zhou, Zuan Gao, Zixiao Wang, Boqiang Zhang et al.CVPR 2025
