Lune

ICLR2026顶会

Synthetic Bootstrapped Pretraining

Zitong Yang, Aonan Zhang, Hong Liu, Tatsunori Hashimoto, Emmanuel J. Candès, Chong Wang, Ruoming Pang

2026年份
3被引次数

摘要

We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dataset and then leverages it to synthesize a vast new corpus for joint training. While the standard pretraining teaches LMs to learn causal correlations among tokens within a single document, it is not designed to efficiently model the rich, learnable inter-document correlations that can potentially lead to better performance. We validate SBP by designing a compute-matched pretraining setup and pretrain a 3B-parameter and a 6B-parameter model on up to 1T tokens from scratch. We find SBP consistently improves upon a strong repetition baseline and delivers up to 60% of performance improvement attainable by an oracle upper bound with access to 20x more unique data. Qualitative analysis reveals that the synthesized documents go beyond mere paraphrases -SBP first abstracts a core concept from the seed material and then crafts a new narration on top of it. Besides strong empirical performance, SBP admits a natural Bayesian interpretation: the synthesizer implicitly learns to abstract the latent concepts shared between related documents. To test our hypothesis, we design a compute-matched, data-constrained experimental framework under which we pretrain a 3B-parameter and a 6B-parameter model on up to 1T tokens from scratch (Li et al., 2024; Zyphra, 2024) , demonstrating the potential applicability of SBP for advancing frontier LMs. We compare SBP's performance against two crucial references: a strong repetition baseline,

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖