Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models
Hong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu Ma
摘要
Language modeling on large-scale datasets leads to impressive performance gains on various downstream language tasks. The validation pre-training loss (or perplexity in autoregressive language modeling) is often used as the evaluation metric when developing language models since the pre-training loss tends to be well-correlated with downstream performance (which is itself difficult to evaluate comprehensively). Contrary to this conventional wisdom, this paper shows that 1) pre-training loss cannot fully explain downstream performance and 2) flatness of the model is well-correlated with downstream performance where pre-training loss is not. On simplified datasets, we identify three ways to produce models with the same (statistically optimal) pre-training loss but different downstream performance: continue pre-training after convergence, increasing the model size, and changing the training algorithm. These experiments demonstrate the existence of implicit bias of pre-training algorithms/optimizers -- among models with the same minimal pre-training loss, they implicitly prefer more transferable ones. Toward understanding this implicit bias, we prove that SGD with standard mini-batch noise implicitly prefers flatter minima in language models, and empirically observe a strong correlation between flatness and downstream performance among models with the same minimal pre-training loss. We also prove in a synthetic language setting that among the models with the minimal pre-training loss, the flattest model transfers to downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- Understanding Emergent Abilities of Language Models from the Loss PerspectiveZhengxiao Du, Aohan Zeng, Yuxiao Dong, Jie TangNeurIPS 2024 · 被引用 113 次
- Improving Language Plasticity via Pretraining with Active ForgettingYihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani 等NeurIPS 2023 · 被引用 49 次
- ModelDiff: A Framework for Comparing Learning AlgorithmsHarshay Shah, Sung Min Park, Andrew Ilyas, Aleksander MadryICML 2023 · 被引用 36 次
- Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsKaiyue Wen, Yuchen Li, Bingbin Liu, Andrej RisteskiNeurIPS 2023 · 被引用 32 次
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi 等ICLR 2026 · 被引用 28 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
相关 Paper
- On the Transferability of Pre-trained Language Models: A Study from Artificial DatasetsDavid Cheng-Han Chiang, Hung-Yi LeeAAAI 2022 · 被引用 33 次
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
- Training Trajectories of Language Models Across ScalesMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin 等ACL 2023 · 被引用 12 次
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase 等ICLR 2026 · 被引用 13 次
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas 等ICLR 2025
