The Best Instruction-Tuning Data are Those That Fit
Dylan Zhang, Qirun Dai, Hao Peng
摘要
High-quality supervised finetuning (SFT) data are essential for unlocking pretrained LLMs' capabilities. Typically, instructions are paired with responses from various sources by humans annotators or other LMs, which are often out of the distribution of the target model to be finetuned. This, at scale, can lead to diminishing returns and even hurt the models' performance and robustness. We hypothesize that SFT is most effective with data aligned to the model's pretrained distribution and propose GRAPE-a novel SFT framework that tailors supervision to the target model. For each instruction, it gathers responses from various sources, and selects the one that aligns most closely to the target model's pretrained distribution, as measured by the normalized probability. We then proceed with standard SFT with these selected responses. We first evaluate GRAPE with a controlled experiment, where we sample various solutions for each question in UltraInteract from multiple models and finetune on GRAPE-selected data using LMs from different families including LLaMA.1-8B, Mistral-7B and Qwen2.5-7B. GRAPE significantly outperforms strong baselines, including distilling from the strongest model with absolute gain up to 13.8% averaging across benchmarks, and a baseline trained on 3× more data with maximum 17.3% performance improvements. GRAPE's strong performance generalizes to off-the-shelf SFT data. We use GRAPE to subsample responses from the post-training data used for Tulu3 and Olmo-2. GRAPE can outperform strong baselines with 4.5 times the data by 6.1% and state-of-the-art data selection approaches by 3.9% on average performance. Remarkably, using 1/3 data and half number of epochs, GRAPE allows LLaMA.1-8B to surpass the performance of Tulu3-SFT by 3.5%. Our findings highlight that aligning supervision with the pretrained distribution offers a simple yet powerful way to improve SFT efficiency and performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 被引用 43 次
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use InsteadFeiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica 等ICLR 2026 · 被引用 27 次
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability ContinuumGaotang Li, Ruizhong Qiu, Xiusi Chen, Heng Ji 等ICML 2026 · 被引用 15 次
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning DistillationKaiyuan Liu, Shaotian Yan, Rui Miao, Bing Wang 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper59
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu 等EMNLP 2025
- Diffusion Instruction TuningChen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe 等ICML 2025
- Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingZhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng 等ICLR 2025
- Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsNing Ding, Yulin Chen, Bokai Xu, Yujia Qin 等EMNLP 2023 · 被引用 95 次
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 被引用 176 次
