Measuring Diversity in Synthetic Datasets
Yuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li, Zibin Zheng, Peilin Zhao, Liang Chen, Yatao Bian
摘要
Large language models (LLMs) are widely adopted to generate synthetic datasets for various natural language processing (NLP) tasks, such as text classification and summarization. However, accurately measuring the diversity of these synthetic datasets-an aspect crucial for robust model performance-remains a significant challenge. In this paper, we introduce DCScore, a novel method for measuring synthetic dataset diversity from a classification perspective. Specifically, DCScore formulates diversity evaluation as a sample classification task, leveraging mutual relationships among samples. We further provide theoretical verification of the diversity-related axioms satisfied by DCScore, highlighting its role as a principled diversity evaluation method. Experimental results on synthetic datasets reveal that DCScore enjoys a stronger correlation with multiple diversity pseudo-truths of evaluated datasets, underscoring its effectiveness. Moreover, both empirical and theoretical evidence demonstrate that DCScore substantially reduces computational costs compared to existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Synthetic Data Generation for Training Diversified Commonsense Reasoning ModelsTianhui Zhang, Bei Peng, Danushka BollegalaACL 2026 · 被引用 1 次
- B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal UnderstandingChangho Choi, Youngwoo Shin, Gyojin Han, Dong-Jae Lee 等ACM MM 2025 · 被引用 1 次
- Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based GuidanceJingwei Zhang, Haoyu LEI, Zijin Feng, Jiacheng Sun 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle 等ICLR 2020 · 被引用 236 次
相关 Paper
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni 等ACL 2026
- Evaluating the Evaluation of Diversity in Commonsense GenerationTianhui Zhang, Bei Peng, Danushka BollegalaACL 2025 · 被引用 6 次
- Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricYuming Yang, Yang Nan, Junjie Ye, Shihan Dou 等ACL 2025 · 被引用 15 次
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
- CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMsSuhas S. Kowshik, Abhishek Divekar, Vijit MalikEMNLP 2024
