Lune

ICML2026顶会

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining

Jie Hao, Rui Yu, Wei Zhang, Huixia Judy Wang, Jie Xu, Mingrui Liu

2026年份
2被引次数

摘要

Effective data selection is essential for pretraining large language models (LLMs), improving efficiency and generalization to downstream tasks. However, existing approaches often rely on external pretrained models, making it difficult to separate the benefits of data selection from those introduced by external models. In addition, many methods estimate data importance from a fixed model state or short-horizon update, making it hard to capture how data preference changes as the model evolves during pretraining. In this paper, we introduce BLISS (BileveL Influence Scoring method for data Selection), a lightweight data selection method that operates entirely from scratch, without external pretrained oracle models, while modeling dynamic data preference. BLISS uses a small proxy model as a surrogate for the LLM and trains a score model to estimate sample importance through multi-step proxy updates induced by score-weighted training data. We formulate data selection as a bilevel optimization problem: the upper-level objective optimizes the score model to assign sample weights, so minimizing the lower-level weighted training loss improves validation performance. Once optimized, the score model predicts influence scores, enabling efficient selection of high-quality samples for LLM pretraining. We validate BLISS by pretraining 410M/1B/2.8B Pythia and LLaMA-0.5B models on selected C4 subsets. Under the 1B setting, BLISS achieves a 1.7×1.7\times speedup in reaching the same performance as the state-of-the-art method, while delivering superior performance across multiple downstream tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖