CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMs
Suhas S. Kowshik, Abhishek Divekar, Vijit Malik
Abstract
Large language models (LLMs) have demonstrated remarkable performance in diverse tasks using zero-shot and few-shot prompting. Even though their capabilities of data synthesis have been studied well in recent years, the generated data suffers from a lack of diversity, less adherence to the prompt, and potential biases that creep into the data from the generator model. In this work, we tackle the challenge of generating datasets with high diversity, upon which a student model is trained for downstream tasks. Taking the route of decoding-time guidance-based approaches, we propose CorrSynth, which generates data that is more diverse and faithful to the input prompt using a correlated sampling strategy. Further, our method overcomes the complexity drawbacks of some other guidance-based techniques like classifier-based guidance. With extensive experiments, we show the effectiveness of our approach and substantiate our claims. In particular, we perform intrinsic evaluation to show the improvements in diversity. Our experiments show that CorrSynth improves both student metrics and intrinsic metrics upon competitive baselines across four datasets, showing the innate advantage of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
Related papers
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource LanguagesTatiana Anikina, Ján Cegin, Jakub Simko, Simon OstermannEMNLP 2025
- SynthesizRR: Generating Diverse Datasets with Retrieval AugmentationAbhishek Divekar, Greg DurrettEMNLP 2024 · 4 citations
- PANGEA: Projection-Based Augmentation with Non-Relevant General Data for Enhanced Domain Adaptation in LLMsSeungyoo Lee, Giung Nam, Moonseok Choi, Hyungi Lee et al.NeurIPS 2025
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni et al.ACL 2026
- Consistency-guided Prompt Learning for Vision-Language ModelsShuvendu Roy, Ali EtemadICLR 2024 · 102 citations
