Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
Ján Cegin, Branislav Pecher, Jakub Simko, Ivan Srba, Mária Bieliková, Peter Brusilovsky
Abstract
The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models. However, more research is needed to assess how different prompts, seed data selection strategies, filtering methods, or model settings affect the quality of paraphrased data (and downstream models). In this study, we investigate three text diversity incentive methods well established in crowdsourcing: taboo words, hints by previous outlier solutions, and chaining on previous outlier solutions. Using these incentive methods as part of instructions to LLMs augmenting text datasets, we measure their effects on generated texts' lexical diversity and downstream model performance. We compare the effects over 5 different LLMs, 6 datasets and 2 downstream models. We show that diversity is most increased by taboo words, but downstream model performance is highest with hints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3df772e1-8acd-4a19-8422-17d1a2e21cf1Cited by top-tier papers4
- Diversity-oriented Data Augmentation with Large Language ModelsZaitian Wang, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu et al.ACL 2025 · 11 citations
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource LanguagesTatiana Anikina, Ján Cegin, Jakub Simko, Simon OstermannEMNLP 2025
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic TriplesKyohoon Jin, Juhwan Choi, Jungmin Yun, Junho Lee et al.EMNLP 2025
Builds on13
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- Task Contamination: Language Models May Not Be Few-Shot AnymoreChangmao Li, Jeffrey FlaniganAAAI 2024 · 138 citations
- Novelty Controlled Paraphrase Generation with Retrieval Augmented Conditional Prompt TuningJishnu Ray Chowdhury, Yong Zhuang, Shuyi WangAAAI 2022 · 39 citations
Related papers
- ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model RobustnessJán Cegin, Jakub Simko, Peter BrusilovskyEMNLP 2023 · 25 citations
- Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt EnhancementPengwei Zhan, Zhen Xu, Qian Tan, Jie Song et al.EMNLP 2024 · 9 citations
- Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborOr Honovich, Thomas Scialom, Omer Levy, Timo SchickACL 2023 · 92 citations
- Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language InferenceArtur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de MarneffeACL 2026
- Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsMinki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong ChoNeurIPS 2024 · 3 citations
