Procedural Pretraining: Warming Up Language Models with Abstract Data
Liangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran, Damien Teney
Abstract
Pretraining language models directly on webscale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to abstract structured data to ease the subsequent acquisition of rich semantic knowledge, much like humans learning simple logic and mathematics before higher reasoning. We focus on procedural data, generated by formal languages and other simple algorithms, as such abstract data. We first diagnose the algorithmic skills that different forms of procedural data can improve, often significantly. For example, the accuracy of context recall (NEEDLE-IN-A-HAYSTACK) jumps from 10 to 98% when a model is pretrained on Dyck sequences (balanced brackets). Second, we study how these gains are reflected in pretraining larger models (up to 1.3B). We find that front-loading as little as 0.1-0.3% procedural data significantly outperforms standard pretraining on natural language, code, and informal mathematics (C4, CODEPARROT, and DEEPMIND-MATH datasets). Notably, this also enables the models to reach the same loss value with only 55 / 67 / 86% of the original data and thus a comparable reduction in FLOPs. Third, we explore the mechanisms behind the benefits and find that procedural pretraining instils non-trivial structure in both attention and MLP layers. The former is particularly important for structured domains (e.g. code), and the latter for language. Finally, we lay a path for combining multiple forms of procedural data. Our results show that procedural pretraining is a simple, lightweight means to improving performance and accelerating language model pretraining, ultimately suggesting the promise of disentangling knowledge acquisition from reasoning in LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 181 citations
Related papers
- On Code-Induced Reasoning in LLMsAbdul Waheed, Zhen Wu, Carolyn Rose, Daphne IppolitoICLR 2026 · 6 citations
- What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure CodeYuze Zhao, Junpeng Fang, Lu Yu, Zhenya Huang et al.ICML 2026 · 1 citation
- When Do Program-of-Thought Works for Reasoning?Zhen Bi, Ningyu Zhang, Yinuo Jiang, Shumin Deng et al.AAAI 2024
- To Code or Not To Code? Exploring Impact of Code in Pre-trainingViraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot et al.ICLR 2025 · 3 citations
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsLaura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara et al.ICLR 2025
