Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
Feiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar, Yuanzhi Zhang, Rongxing Du, Anit Kumar Sahu, Ruoxi Jia
摘要
This work focuses on leveraging and selecting from vast, unlabeled, open data to pre-fine-tune a pre-trained language model. The goal is to minimize the need for costly domain-specific data for subsequent fine-tuning while achieving desired performance levels. While many data selection algorithms have been designed for small-scale applications, rendering them unsuitable for our context, some emerging methods do cater to language data scales. However, they often prioritize data that aligns with the target distribution. While this strategy may be effective when training a model from scratch, it can yield limited results when the model has already been pre-trained on a different distribution. Differing from prior work, our key idea is to select data that nudges the pre-training distribution closer to the target distribution. We show the optimality of this approach for fine-tuning tasks under certain conditions. We demonstrate the efficacy of our methodology across a diverse array of tasks (NLU, NLG, zero-shot) with models up to 2.7B, showing that it consistently surpasses other selection methods. Moreover, our proposed method is significantly faster than existing techniques, scaling to millions of samples within a single GPU hour. Our code is open-sourced (Code repository: https://anonymous.4open.science/r/DV4LLM-D761/ ). While fine-tuning offers significant potential for enhancing performance across diverse tasks, its associated costs often limit its widespread adoption; with this work, we hope to lay the groundwork for cost-effective fine-tuning, making its benefits more accessible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Not All Tokens Are What You Need for PretrainingZhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 等NeurIPS 2024 · 被引用 99 次
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout ReplayYifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen 等NeurIPS 2025 · 被引用 63 次
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 被引用 59 次
- A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao 等ACL 2025 · 被引用 12 次
- Safety at One Shot: Patching Fine-Tuned LLMs with A Single InstanceJiawen Zhang, Lipeng He, Kejia Chen, Jian Lou 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Selecting Large Language Model to Fine-tune via Rectified Scaling LawHaowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen 等ICML 2024 · 被引用 32 次
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding 等ICML 2025
- Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction TuningBardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel 等CVPR 2025
- Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task GeneralizationYuxian Gu, Pei Ke, Xiaoyan Zhu, Minlie HuangEMNLP 2022 · 被引用 3 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
