Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
Michael Y. Hu, Jackson Petty, Chuan Shi, William Merrill, Tal Linzen
Abstract
Pretraining language models on formal language can improve their acquisition of natural language. Which features of the formal language impart an inductive bias that leads to effective transfer? Drawing on insights from linguistics and complexity theory, we hypothesize that effective transfer occurs when two conditions are met: the formal language should capture the dependency structures present in natural language, and it should remain within the computational limitations of the model architecture. We experiment with pre-pretraining (training on formal language before natural languages) on transformers and find that formal languages capturing hierarchical dependencies indeed enable language models to achieve lower loss on natural language and better linguistic generalization compared to other formal languages. We also find modest support for the hypothesis that the formal language should fall within the computational limitations of the architecture. Strikingly, pre-pretraining reduces loss more efficiently than training on a matched amount of natural language. For a 1B-parameter language model trained on roughly 1.6B tokens of natural language, pre-pretraining achieves the same loss and better linguistic generalization with a 33% smaller token budget. Finally, we also give mechanistic evidence of transfer from formal to natural language: attention heads acquired during pre-pretraining remain crucial for the model's performance on syntactic evaluations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ee89cd2-90b2-4b90-a6dd-c4ba05647c70Cited by top-tier papers10
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou et al.ICLR 2026 · 29 citations
- Can You Learn to See Without Images? Procedural Warm-Up for Vision TransformersZachary Shinnick, Liangze Jiang, Hemanth Saratchandran, Damien Teney et al.CVPR 2026 · 9 citations
- Procedural Pretraining: Warming Up Language Models with Abstract DataLiangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran et al.ICML 2026 · 6 citations
- Unraveling Syntax: Language Modeling and the Substructure of GrammarsLaura Ying Schulz, Daniel Mitropolsky, Tomaso A PoggioICML 2026 · 5 citations
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient ReasonersXin Xu, Clive Bai, Kai Yang, Tianhao Chen et al.ICLR 2026 · 5 citations
Builds on27
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao et al.ICML 2023 · 287 citations
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 236 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
Related papers
- How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive BiasesAaron Mueller, Tal LinzenACL 2023 · 9 citations
- Pretraining with Artificial Language: Studying Transferable Knowledge in Language ModelsRyokan Ri, Yoshimasa TsuruokaACL 2022 · 40 citations
- SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by SimulationMatthias Lindemann, Alexander Koller, Ivan TitovACL 2024 · 2 citations
- On the Transferability of Pre-trained Language Models: A Study from Artificial DatasetsDavid Cheng-Han Chiang, Hung-Yi LeeAAAI 2022 · 33 citations
- Language Acquisition Device in Large Language ModelsMasato Mita, Taiga Someya, Ryo Yoshida, Yohei OsekiACL 2026
