How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive Biases
Aaron Mueller, Tal Linzen
Abstract
Accurate syntactic representations are essential for robust generalization in natural language. Recent work has found that pre-training can teach language models to rely on hierarchical syntactic features—as opposed to incorrect linear features—when performing tasks after fine-tuning. We test what aspects of pre-training are important for endowing encoder-decoder Transformers with an inductive bias that favors hierarchical syntactic generalizations. We focus on architectural features (depth, width, and number of parameters), as well as the genre and size of the pre-training corpus, diagnosing inductive biases using two syntactic transformation tasks: question formation and passivization, both in English. We find that the number of parameters alone does not explain hierarchical generalization: model depth plays greater role than model width. We also find that pre-training on simpler language, such as child-directed speech, induces a hierarchical bias using an order-of-magnitude less data than pre-training on more typical datasets based on web text or Wikipedia; this suggests that in cognitively plausible language acquisition settings, neural language models may be more data-efficient than previously thought.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c13b43d2-dd32-43c8-8bd0-01a2342a1a89Cited by top-tier papers12
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 79 citations
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- Random Scaling of Emergent CapabilitiesRosie Zhao, Tian Qin, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 3 citations
- Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across TasksDamien Teney, Liangze Jiang, Hemanth Saratchandran, Simon LuceyICLR 2026 · 3 citations
- StrAE: Autoencoding for Pre-Trained Embeddings using Explicit StructureMattia Opper, Victor Prokhorov, Siddharth NarayanaswamyEMNLP 2023 · 1 citation
Builds on8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Predicting Inductive Biases of Pre-Trained ModelsCharles Lovering, Rohan Jha, Tal Linzen, Ellie PavlickICLR 2021 · 70 citations
- Scale Efficiently: Insights from Pretraining and Finetuning TransformersYi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus et al.ICLR 2022 · 67 citations
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 40 citations
Related papers
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 17 citations
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill et al.ACL 2025
- Semantic Training Signals Promote Hierarchical Syntactic Generalization in TransformersAditya Yedetore, Najoung KimEMNLP 2024 · 4 citations
- Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic TransformationsMatthias Lindemann, Alexander Koller, Ivan TitovEMNLP 2024
- SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by SimulationMatthias Lindemann, Alexander Koller, Ivan TitovACL 2024 · 2 citations
