Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions
John Joon Young Chung, Ece Kamar, Saleema Amershi
摘要
Large language models (LLMs) can be used to generate text data for training and evaluating other models. However, creating high-quality datasets with LLMs can be challenging. In this work, we explore human-AI partnerships to facilitate high diversity and accuracy in LLM-based text data generation. We first examine two approaches to diversify text generation: 1) logit suppression, which minimizes the generation of languages that have already been frequently generated, and 2) temperature sampling, which flattens the token sampling probability. We found that diversification approaches can increase data diversity but often at the cost of data accuracy (i.e., text and labels being appropriate for the target domain). To address this issue, we examined two human interventions, 1) label replacement (LR), correcting misaligned labels, and 2) out-of-scope filtering (OOSF), removing instances that are out of the user's domain of interest or to which no considered label applies. With oracle studies, we found that LR increases the absolute accuracy of models trained with diversified datasets by 14.4%. Moreover, we found that some models trained with data generated with LR interventions outperformed LLM-based few-shot classification. In contrast, OOSF was not effective in increasing model accuracy, implying the need for future work in human-in-the-loop text data generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 被引用 102 次
- Curiosity-driven Red-teaming for Large Language ModelsZhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang 等ICLR 2024 · 被引用 84 次
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith 等ICLR 2026 · 被引用 41 次
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh 等EMNLP 2023 · 被引用 30 次
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
它引用的顶会 Paper10
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros 等CVPR 2022 · 被引用 198 次
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 被引用 99 次
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian 等ACL 2022 · 被引用 91 次
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma 等EMNLP 2020 · 被引用 62 次
相关 Paper
- CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMsSuhas S. Kowshik, Abhishek Divekar, Vijit MalikEMNLP 2024
- Generating Diverse Training Samples for Relation Extraction with Large Language ModelsZexuan Li, Hongliang Dai, Piji LiACL 2025
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationGaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. LaradjiEMNLP 2023 · 被引用 11 次
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource LanguagesTatiana Anikina, Ján Cegin, Jakub Simko, Simon OstermannEMNLP 2025
