Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
Yuming Yang, Yang Nan, Junjie Ye, Shihan Dou, Xiao Wang, Shuo Li, Huijie Lv, Tao Gui, Qi Zhang, Xuanjing Huang
Abstract
Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive finetuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level "novelty." Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric. The code is available at https: //github.com/UmeanNever/NovelSum .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c04cedcf-ce20-4fa7-8ace-5faec97318c7Cited by top-tier papers4
- Optimizing Diversity and Quality through Base-Aligned Model CollaborationYichen Wang, Chenghao Yang, Tenghao Huang, Muhao Chen et al.ICML 2026 · 7 citations
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM ReasoningWeize Liu, Yongchi Zhao, Yijia Luo, Mingyu Xu et al.ICLR 2026 · 6 citations
- SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model TrainingPowei Chang, Jinpeng Zhang, Bowen Chen, Chenyu Wang et al.ICLR 2026 · 5 citations
- Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan et al.ACL 2026 · 5 citations
Builds on11
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction TuningWei Liu, Weihao Zeng, Keqing He, Yong Jiang et al.ICLR 2024 · 369 citations
Related papers
- Measuring Diversity in Synthetic DatasetsYuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li et al.ICML 2025
- Priority on High-Quality: Selecting Instruction Data via Consistency Verification of Noise InjectionHong Zhang, Feng Zhao, Ruilin Zhao, Cheng Yan et al.EMNLP 2025
- From Macro to Micro: Probing Dataset Diversity in Language Model Fine-TuningHaoyu Li, Xuhong Li, Yiming Dong, Kun LiuAAAI 2026 · 2 citations
- The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite GraphMinghao Wu, Thuy-Trang Vu, Lizhen Qu, Gholamreza HaffariICML 2025
- G-DIG: Towards Gradient-based DIverse and hiGh-quality Instruction Data Selection for Machine TranslationXingyuan Pan, Luyang Huang, Liyan Kang, Zhicheng Liu et al.ACL 2024
