The Coverage Principle: How Pre-Training Enables Post-Training
Fan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi, Adam Block, Jordan T. Ash, Akshay Krishnamurthy, Dylan J Foster
摘要
Language models demonstrate remarkable abilities when pre-trained on large text corpora and fine-tuned for specific tasks, but how and why pre-training shapes the success of the final model remains poorly understood. Notably, although pre-training success is often quantified by cross entropy loss, cross entropy can be poorly predictive of downstream performance. Instead, we provide a theoretical perspective on this relationship through the lens of coverage, which quantifies the probability mass the pre-trained model places on high-quality responses and which is necessary and sufficient for post-training and test-time scaling methods like Best-of-N to succeed. Our main results develop an understanding of the coverage principle, a phenomenon whereby next-token prediction implicitly optimizes toward a model with good coverage. In particular, we uncover a mechanism that explains the power of coverage in predicting downstream performance: coverage generalizes faster than cross entropy, avoiding spurious dependence on problem dependent parameters such as the sequence length. We also study practical algorithmic interventions with provable benefits for improving coverage, including (i) model/checkpoint selection procedures, (ii) gradient normalization schemes, and (iii) test-time decoding strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningAndrew Wagenmaker, Perry Dong, Raymond Tsao, Chelsea Finn 等ICML 2026 · 被引用 10 次
- From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL FinetuningZhanyi Sun, shuran songICML 2026 · 被引用 6 次
- On the Emergence of Implicit Curriculum in RLVR Learning DynamicsYu Huang, Zixin Wen, Yuejie Chi, Yuting Wei 等ICML 2026 · 被引用 6 次
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model ReasoningLingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma 等ICML 2026 · 被引用 1 次
- A Theoretical Framework for Statistical Evaluability of Generative ModelsShashaank Aiyer, Yishay Mansour, Shay Moran, Han ShaoICML 2026 · 被引用 1 次
它引用的顶会 Paper31
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan 等NeurIPS 2023 · 被引用 420 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 被引用 194 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 被引用 144 次
相关 Paper
- Improving Diversity in Language Models: When Temperature Fails, Change the LossAlexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen 等ICML 2025
- Temporal Scaling Law for Large Language ModelsYizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen 等EMNLP 2025
- A Mathematical Exploration of Why Language Models Help Solve Downstream TasksNikunj Saunshi, Sadhika Malladi, Sanjeev AroraICLR 2021 · 被引用 93 次
- A Simple Model of Inference Scaling LawsNoam Itzhak LeviICML 2025
- Large Language Models to Diffusion FinetuningEdoardo Cetin, Tianyu Zhao, Yujin TangICML 2025
