SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
Yexiao He, Ziyao Wang, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang, Ang Li
Abstract
The pre-trained Large Language Models (LLMs) can be adapted for many downstream tasks and tailored to align with human preferences through fine-tuning. Recent studies have discovered that LLMs can achieve desirable performance with only a small amount of high-quality data, suggesting that a large amount of the data in these extensive datasets is redundant or even harmful. Identifying high-quality data from vast datasets to curate small yet effective datasets has emerged as a critical challenge. In this paper, we introduce SHED, an automated dataset refinement framework based on Shapley value for instruction fine-tuning. SHED eliminates the need for human intervention or the use of commercial LLMs. Moreover, the datasets curated through SHED exhibit transferability, indicating they can be reused across different LLMs with consistently high performance. We conduct extensive experiments to evaluate the datasets curated by SHED. The results demonstrate SHED's superiority over state-of-the-art methods across various tasks and LLMs; notably, datasets comprising only 10% of the original data selected by SHED achieve performance comparable to or surpassing that of the full datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24b73f32-0d5a-4538-a283-e4730edbc7dfCited by top-tier papers7
- T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction TuningYanjun Fu, Faisal Hamman, Sanghamitra DuttaNeurIPS 2025 · 15 citations
- Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal ModelsXu Yang, Chen Liu, Ying WeiNeurIPS 2025 · 2 citations
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang et al.AAAI 2026 · 1 citation
- Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM FinetuningHao Cheng, Rui Zhang, Ling Li, Na Di et al.ICLR 2026
- Adalina: Adaptive Linear Approximation for the Shapley Value and BeyondWeida Li, Yaoliang Yu, Bryan Kian Hsiang LowICML 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
Related papers
- From Selection to Refinement: Iterative Optimization for Instruction DataHang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou et al.ACL 2026
- Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessQifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu et al.ICCV 2025 · 1 citation
- Automatic Instruction Evolving for Large Language ModelsWeihao Zeng, Can Xu, Yingxiu Zhao, Jian-Guang Lou et al.EMNLP 2024 · 2 citations
- Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleFan Zhou, Zengzhi Wang, Qian Liu, Junlong Li et al.ICML 2025
- Exploring Parameter-Efficient Fine-Tuning of Large Language Model on Automated Program RepairGuochang Li, Chen Zhi, Jialiang Chen, Junxiao Han et al.ASE 2024 · 8 citations
