Train Once, Answer All: Many Pretraining Experiments for the Cost of One
Sebastian Bordt, Martin Pawelczyk
Abstract
Recent work has demonstrated that controlled pretraining experiments are a powerful tool for studying the relationship between training data and large language model (LLM) behavior. However, the computational cost of pretraining presents a significant constraint. To overcome this constraint, we propose a new approach where multiple experiments are conducted simultaneously during a single training run. We validate our approach by performing ten experiments while training on 210B tokens, with models of up to 2.7B parameters. Although models are trained only once, we can replicate the results of multiple previous works on data contamination, poisoning, and memorization. We also conduct novel investigations into knowledge acquisition, mathematical reasoning, and watermarking. For example, we dynamically update the training data until a model acquires a particular piece of knowledge. Remarkably, the influence of the experiments on the model's training dynamics and overall performance is minimal. However, interactions between experiments may act as a confounder in our approach. We propose continual pretraining dependence testing (CPDT), a novel technique to test for interactions with continual pretraining experiments, finding them to be negligible in our setup. Overall, our results suggest that performing multiple pretraining experiments within a single training run can enable rigorous scientific experimentation with large models on a compute budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af4dae17-0761-4f44-90d3-ff5ef14203b1Cited by top-tier papers1
Ask how each one uses itBuilds on46
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li et al.ACL 2026
- Persistent Pre-training Poisoning of LLMsYiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi et al.ICLR 2025
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang et al.NeurIPS 2024 · 124 citations
- Optimal Splitting of Language Models from Mixtures to Specialized DomainsSkyler Seto, Pierre Ablin, Anastasiia Filippova, Jiayuan Ye et al.ICML 2026 · 2 citations
- Reusing Pre-Training Data at Test Time is a Compute MultiplierAlex Fang, Thomas Voice, Ruoming Pang, Ludwig Schmidt et al.ICLR 2026 · 4 citations
