Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming Yin
摘要
The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach. However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks. To better understand factors that moderate the effectiveness of the LLMgenerated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification. Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data. We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Is In-Context Learning in Large Language Models Bayesian? A Martingale PerspectiveFabian Falck, Ziyu Wang, Christopher C. HolmesICML 2024 · 被引用 46 次
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning 等NSDI 2026 · 被引用 29 次
- WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive StorytellingZhuoran Lu, Qian Zhou, Yi WangCHI 2025 · 被引用 25 次
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao 等NeurIPS 2025 · 被引用 17 次
- Exploring Empty Spaces: Human-in-the-Loop Data AugmentationCatherine Yeh, Donghao Ren, Yannick Assogba, Dominik Moritz 等CHI 2025 · 被引用 13 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human SolutionsJiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker 等CHI 2023 · 被引用 283 次
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 被引用 244 次
相关 Paper
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
- Data-Constrained Synthesis of Training Data for De-IdentificationThomas Vakili, Aron Henriksson, Hercules DalianisACL 2025 · 被引用 3 次
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen 等EMNLP 2024 · 被引用 2 次
- Evaluating Language Models as Synthetic Data GeneratorsSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan 等ACL 2025
- Synthetic Text Generation for Training Large Language Models via Gradient MatchingDang Nguyen, Zeman Li, MohammadHossein Bateni, Vahab Mirrokni 等ICML 2025
