PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
Oskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra, Hailey Schoelkopf, Willem H. Zuidema, Stella Biderman
Abstract
The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly different results in response to slight variations in initial conditions, e.g., the random seed. Crucially, the research community still lacks sufficient resources and tools to systematically investigate pre-training stability, particularly for decoder-only language models. We introduce the PolyPythias, a set of 45 new training runs for the Pythia model suite: 9 new seeds across 5 model sizes, from 14M to 410M parameters, resulting in about 7k new checkpoints that we release. Using these new 45 training runs, in addition to the 5 already available, we study the effects of different initial conditions determined by the seed-i.e., parameters' initialisation and data order-on (i) downstream performance, (ii) learned linguistic representations, and (iii) emergence of training phases. In addition to common scaling behaviours, our analyses generally reveal highly consistent training dynamics across both model sizes and initial conditions. Further, the new seeds for each model allow us to identify outlier training runs and delineate their characteristics. Our findings show the potential of using these methods to predict training stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a56a86cc-92f9-4c81-9259-528c0623a83eCited by top-tier papers7
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim et al.ICML 2026 · 22 citations
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and ScaleJames A. Michaelov, Roger P. Levy, Benjamin BergenNeurIPS 2025 · 15 citations
- Evolution of Concepts in Language Model Pre-TrainingXuyang Ge, Wentao Shu, Jiaxing Wu, Yunhua Zhou et al.ICLR 2026 · 8 citations
- The Mechanistic Emergence of Symbol Grounding in Language ModelsShuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang et al.ICML 2026 · 4 citations
- Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov et al.ACL 2026 · 2 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
Related papers
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Tracing the Representation Geometry of Language Models from Pretraining to Post-trainingMelody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru et al.NeurIPS 2025 · 38 citations
- Evaluating n-Gram Novelty of Language Models Using Rusty-DAWGWilliam Merrill, Noah A. Smith, Yanai ElazarEMNLP 2024 · 2 citations
- Training Trajectories of Language Models Across ScalesMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin et al.ACL 2023 · 12 citations
- CausalGym: Benchmarking causal interpretability methods on linguistic tasksAryaman Arora, Dan Jurafsky, Christopher PottsACL 2024 · 3 citations
