DavIR: Data Selection via Implicit Reward for Large Language Models
Haotian Zhou, Tingkai Liu, Qianli Ma, Yufeng Zhang, Jianbo Yuan, Pengfei Liu, Yang You, Hongxia Yang
Abstract
We introduce DavIR, a model-based data selection method for post-training Large Language Models. DavIR generalizes Reducible Holdout Loss to core-set selection problem of causal language modeling, and quantifies the "learnability" of a given datum with respect to a pre-trained LLM based on relative reduction in loss during fine-tuning, a metric we show to be closely related to the implicit reward model described in Direct Preference Optimization (DPO). We show that 6% of Alpaca dataset selected with DavIR can steer both the LLaMA and Gemma model family to produce superior performance compared to the same models trained on the full 52K dataset. We also show that Alpaca dataset compressed with DavIR can be combined with GSM8K dataset to effectively balance open-domain freeform QA and mathematical reasoning capabilities. Finally, we apply the DavIR objective to DPO and develop a normalized DavIR-DPO objective which improves alignment performance of Zephyr-7B-SFT model by 8% (relative) on Al-pacaEval, compared against training on vanilla DPO objective. * Corresponding authors 2023). Selecting the most effective training data during this stage is particularly important since effective steering of LLM during SFT could be achieved by just a few thousand carefully curated data (Zhou et al., 2023) . Previous approaches to selecting SFT training data focused on data quality and diversity (Ji et al., 2023; Zhou et al., 2023; Chen et al., 2023b,a; Li et al., 2023a), guided by the intuition of encouraging LLMs to output accurate and reliable information while maintaining generalization capabilities to a wide range of tasks and scenarios. However, by focusing on the quality and diversity of the data, existing methods are data-centric, and are agnostic to the capabilities of the pretrained model upon which fine-tuning occurs. Instead, following the "Superficial Alignment Hypothesis" (Zhou et al., 2023) which postulates that fine-tuning process unlocks the capabilities of pre-trained LLMs, we seek a model-centric data selection algorithm that chooses data that:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 15 citations
- Reasoning Quality Emerges Early: Data Curation for Reasoning ModelsHongyi Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato et al.ICML 2026
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop LearningZinan Tang, Xin Gao, Qizhi Pei, Zhuoshi Pan et al.EMNLP 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- Normalized Rewards for Preference OptimizationShawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald et al.ICML 2026 · 571 citations
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao et al.AAAI 2025 · 7 citations
- Select Before Use: On the Importance of Reference Model Selection in Preference AlignmentMuyang Li, Runze Wu, Xiangyu Zhao, Bo Han et al.ACL 2026
- Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsBo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao et al.NeurIPS 2025 · 23 citations
- Principled Data Selection for Alignment: The Hidden Risks of Difficult ExamplesChengqian Gao, Haonan Li, Liu Liu, Zeke Xie et al.ICML 2025
