Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
Yang Zhao, Li Du, Xiao Ding, Yangou Ouyang, Hepeng Wang, Kai Xiong, Jinglong Gao, Zhouhao Sun, Dongliang Xu, Qing Yang, Dongchen Li, Bing Qin, Ting Liu
Abstract
Large language models (LLMs) have shown great potential across various industries due to their remarkable ability to generalize through instruction tuning. However, the limited availability of domain-specific data significantly hampers their performance on specialized tasks. While existing methods primarily focus on selecting training data from general datasets that are similar to the target domain, they often fail to consider the joint distribution of instructions, resulting in inefficient learning and suboptimal knowledge transfer. To address these challenges, we introduce G2IS (Gradient-based Graph Instruction Selection), a novel method that constructs a mixed gradient-based instruction graph to capture the joint distribution and interdependencies between instructions. By accounting for the relationships between instructions, G2IS improves domain adaptation efficiency. Additionally, we propose a gradient walk algorithm to refine the data selection process, enhancing both training effectiveness and efficiency. Our experiments demonstrate that G2IS outperforms traditional methods across various domain adaptation tasks, yielding significant performance gains, particularly in complex, data-scarce scenarios. These results underscore the potential of G2IS in advancing the development of large, domain-specific models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 514a7a6d-fabd-414c-af7f-b6cd4361a573Cited by top-tier papers4
- UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data SelectionYang Zhao, Kai Xiong, Xiao Ding, Li Du et al.NeurIPS 2025 · 4 citations
- Computational Budget Should Be Considered in Data SelectionWeilin Wan, Weizhong Zhang, Cheng JinNeurIPS 2025 · 2 citations
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang et al.AAAI 2026 · 1 citation
- LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction TuningYangfan Ye, Xiaocheng Feng, Xiachong Feng, Lei Huang et al.AAAI 2026
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM FinetuningHao Cheng, Rui Zhang, Ling Li, Na Di et al.ICLR 2026
- Automatic Instruction Data Selection for Large Language Models via Uncertainty-Aware Influence MaximizationJindong Han, Hao Liu, Jun Fang, Naiqiang Tan et al.WWW 2025 · 5 citations
- Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationFanqi Wan, Xinting Huang, Tao Yang, Xiaojun Quan et al.EMNLP 2023 · 2 citations
- A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)Nihal Nayak, Paula Rodriguez-Diaz, Neha Hulkund, Sara Beery et al.ICML 2026 · 2 citations
