Lune

ACL2025顶会

DavIR: Data Selection via Implicit Reward for Large Language Models

Haotian Zhou, Tingkai Liu, Qianli Ma, Yufeng Zhang, Jianbo Yuan, Pengfei Liu, Yang You, Hongxia Yang

2025年份
7被引次数
4顶会引用

摘要

We introduce DavIR, a model-based data selection method for post-training Large Language Models. DavIR generalizes Reducible Holdout Loss to core-set selection problem of causal language modeling, and quantifies the "learnability" of a given datum with respect to a pre-trained LLM based on relative reduction in loss during fine-tuning, a metric we show to be closely related to the implicit reward model described in Direct Preference Optimization (DPO). We show that 6% of Alpaca dataset selected with DavIR can steer both the LLaMA and Gemma model family to produce superior performance compared to the same models trained on the full 52K dataset. We also show that Alpaca dataset compressed with DavIR can be combined with GSM8K dataset to effectively balance open-domain freeform QA and mathematical reasoning capabilities. Finally, we apply the DavIR objective to DPO and develop a normalized DavIR-DPO objective which improves alignment performance of Zephyr-7B-SFT model by 8% (relative) on Al-pacaEval, compared against training on vanilla DPO objective. * Corresponding authors 2023). Selecting the most effective training data during this stage is particularly important since effective steering of LLM during SFT could be achieved by just a few thousand carefully curated data (Zhou et al., 2023) . Previous approaches to selecting SFT training data focused on data quality and diversity (Ji et al., 2023; Zhou et al., 2023; Chen et al., 2023b,a; Li et al., 2023a), guided by the intuition of encouraging LLMs to output accurate and reliable information while maintaining generalization capabilities to a wide range of tasks and scenarios. However, by focusing on the quality and diversity of the data, existing methods are data-centric, and are agnostic to the capabilities of the pretrained model upon which fine-tuning occurs. Instead, following the "Superficial Alignment Hypothesis" (Zhou et al., 2023) which postulates that fine-tuning process unlocks the capabilities of pre-trained LLMs, we seek a model-centric data selection algorithm that chooses data that:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖