Measuring Diversity in Synthetic Datasets
Yuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li, Zibin Zheng, Peilin Zhao, Liang Chen, Yatao Bian
Abstract
Large language models (LLMs) are widely adopted to generate synthetic datasets for various natural language processing (NLP) tasks, such as text classification and summarization. However, accurately measuring the diversity of these synthetic datasets-an aspect crucial for robust model performance-remains a significant challenge. In this paper, we introduce DCScore, a novel method for measuring synthetic dataset diversity from a classification perspective. Specifically, DCScore formulates diversity evaluation as a sample classification task, leveraging mutual relationships among samples. We further provide theoretical verification of the diversity-related axioms satisfied by DCScore, highlighting its role as a principled diversity evaluation method. Experimental results on synthetic datasets reveal that DCScore enjoys a stronger correlation with multiple diversity pseudo-truths of evaluated datasets, underscoring its effectiveness. Moreover, both empirical and theoretical evidence demonstrate that DCScore substantially reduces computational costs compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fa78b46-2e67-4f28-9904-b6e30c81157bCited by top-tier papers3
- Synthetic Data Generation for Training Diversified Commonsense Reasoning ModelsTianhui Zhang, Bei Peng, Danushka BollegalaACL 2026 · 1 citation
- B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal UnderstandingChangho Choi, Youngwoo Shin, Gyojin Han, Dong-Jae Lee et al.ACM MM 2025 · 1 citation
- Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based GuidanceJingwei Zhang, Haoyu LEI, Zijin Feng, Jiacheng Sun et al.ICML 2026 · 1 citation
Builds on14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi et al.ICML 2020 · 553 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
Related papers
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni et al.ACL 2026
- Evaluating the Evaluation of Diversity in Commonsense GenerationTianhui Zhang, Bei Peng, Danushka BollegalaACL 2025 · 6 citations
- Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricYuming Yang, Yang Nan, Junjie Ye, Shihan Dou et al.ACL 2025 · 15 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMsSuhas S. Kowshik, Abhishek Divekar, Vijit MalikEMNLP 2024
