Diversity-oriented Data Augmentation with Large Language Models
Zaitian Wang, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu, Pengfei Wang, Yuanchun Zhou
Abstract
Data augmentation is an essential technique in natural language processing (NLP) for enriching training datasets by generating diverse samples. This process is crucial for improving the robustness and generalization capabilities of NLP models. However, a significant challenge remains: Insufficient Attention to Sample Distribution Diversity. Most existing methods focus on increasing the sample numbers while neglecting the sample distribution diversity, which can lead to model overfitting. In response, we explore data augmentation's impact on dataset diversity and propose a Diversity-oriented data Augmentation framework (DoAug). % Specifically, we utilize a diversity-oriented fine-tuning approach to train an LLM as a diverse paraphraser, which is capable of augmenting textual datasets by generating diversified paraphrases. Then, we apply the LLM paraphraser to a selected coreset of highly informative samples and integrate the paraphrases with the original data to create a more diverse augmented dataset. Finally, we conduct extensive experiments on 12 real-world textual datasets. The results show that our fine-tuned LLM augmenter improves diversity while preserving label consistency, thereby enhancing the robustness and performance of downstream tasks. Specifically, it achieves an average performance gain of , surpassing the runner-up baseline with more than three percentage points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ebe6178-4b53-4fa5-b85b-dfd358d2b469Cited by top-tier papers9
- ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense RetrievalFengran Mo, Jinghan Zhang, Yuchen Hui, Jia Ao Sun et al.AAAI 2026 · 7 citations
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 2 citations
- Synthetic Data Generation for Training Diversified Commonsense Reasoning ModelsTianhui Zhang, Bei Peng, Danushka BollegalaACL 2026 · 1 citation
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- Exploring Quality and Diversity in Synthetic Data Generation for Argument MiningJianzhu Bao, Yuqi Huang, Yang Sun, Wenya Wang et al.EMNLP 2025
- Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentationJán Cegin, Branislav Pecher, Jakub Simko, Ivan Srba et al.ACL 2024 · 5 citations
- Paraphrase Augmented Task-Oriented Dialog GenerationSilin Gao, Yichi Zhang, Zhijian Ou, Zhou YuACL 2020 · 78 citations
- Dually Self-Improved Counterfactual Data Augmentation Using Large Language ModelLuhao Zhang, Xinyu Zhang, Linmei Hu, Dandan Song et al.ACL 2025 · 1 citation
- Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsMinki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong ChoNeurIPS 2024 · 3 citations
