Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming Yin
Abstract
The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach. However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks. To better understand factors that moderate the effectiveness of the LLMgenerated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification. Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data. We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2854b30d-de55-4e58-856c-45b6058ff70dCited by top-tier papers42
- Is In-Context Learning in Large Language Models Bayesian? A Martingale PerspectiveFabian Falck, Ziyu Wang, Christopher C. HolmesICML 2024 · 46 citations
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
- WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive StorytellingZhuoran Lu, Qian Zhou, Yi WangCHI 2025 · 25 citations
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao et al.NeurIPS 2025 · 17 citations
- Exploring Empty Spaces: Human-in-the-Loop Data AugmentationCatherine Yeh, Donghao Ren, Yannick Assogba, Dominik Moritz et al.CHI 2025 · 13 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human SolutionsJiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker et al.CHI 2023 · 283 citations
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 244 citations
Related papers
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma et al.ICLR 2025
- Data-Constrained Synthesis of Training Data for De-IdentificationThomas Vakili, Aron Henriksson, Hercules DalianisACL 2025 · 3 citations
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen et al.EMNLP 2024 · 2 citations
- Evaluating Language Models as Synthetic Data GeneratorsSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan et al.ACL 2025
- Synthetic Text Generation for Training Large Language Models via Gradient MatchingDang Nguyen, Zeman Li, MohammadHossein Bateni, Vahab Mirrokni et al.ICML 2025
