Lune

ICML2026顶会

RePro: Training Language Models to Faithfully Recycle the Web for Pretraining

Zichun Yu, Chenyan Xiong

2026年份
2被引次数

摘要

High-quality pretraining data is the fossil fuel of large language models (LLMs), yet its reserves are running low for frontier models. In this paper, we introduce REPRO, a novel web recycling method that trains a relatively small LM with reinforcement learning to generate effective and faithful rephrasings of pretraining data. Specifically, we design one quality reward and three faithfulness rewards, optimizing the LM rephraser to convert organic data into high-quality rephrasings while maintaining its core semantics and structure. In our experiment, we train a 4B rephraser to recycle 72B tokens sampled from DCLM-RefinedWeb. Pretraining results on 400M and 1.4B models demonstrate that REPRO delivers 4.7%-14.0% relative accuracy gains over organic-only baseline on 22 downstream tasks. REPRO also outperforms ReWire, the state-of-the-art web recycling method that prompts a 70B rephraser, as well as the organic baseline with a 4× larger data pool. Experiments with different amounts of recycled data highlight that REPRO improves organic data efficiency by 2-3×. Individual and distributional analyses validate that REPRO preserves more critical information and faithfully reflects the characteristics of organic data compared to prompting-based methods. Together, these results show that REPRO provides an efficient and controllable path to effectively harness the "fossil fuel" of LLM pretraining. We open-source our code, rephraser, and recycled data at https://github.com/cxcscmu/RePro .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖