Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
Matthias Lindemann, Alexander Koller, Ivan Titov
Abstract
Models need appropriate inductive biases to effectively learn from small amounts of data and generalize systematically outside of the training distribution. While Transformers are highly versatile and powerful, they can still benefit from enhanced structural inductive biases for seq2seq tasks, especially those involving syntactic transformations, such as converting active to passive voice or semantic parsing. In this paper, we propose to strengthen the structural inductive bias of a Transformer by intermediate pre-training to perform synthetically generated syntactic transformations of dependency trees given a description of the transformation. Our experiments confirm that this helps with fewshot learning of syntactic tasks such as chunking, and also improves structural generalization for semantic parsing. Our analysis shows that the intermediate pre-training leads to attention heads that keep track of which syntactic transformation needs to be applied to which token, and that the model can leverage these attention heads on downstream tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb6a71d1-887e-492f-8574-bcd750a302e5Cited by top-tier papers2
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- Learning syntax without semantics: Disentangled tiny language modelsEzra Winston, Zico KolterICML 2026
Builds on13
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Improving AMR Parsing with Sequence-to-Sequence Pre-trainingDongqin Xu, Junhui Li, Muhua Zhu, Min Zhang et al.EMNLP 2020 · 57 citations
- Sequence-to-Sequence Learning with Latent Neural GrammarsYoon KimNeurIPS 2021 · 44 citations
- Benchmarking Meaning Representations in Neural Semantic ParsingJiaqi Guo, Qian Liu, Jian-Guang Lou, Zhenwen Li et al.EMNLP 2020 · 24 citations
Related papers
- SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by SimulationMatthias Lindemann, Alexander Koller, Ivan TitovACL 2024 · 2 citations
- How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive BiasesAaron Mueller, Tal LinzenACL 2023 · 9 citations
- Retrofitting Structure-aware Transformer Language Model for End TasksHao Fei, Yafeng Ren, Donghong JiEMNLP 2020 · 56 citations
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill et al.ACL 2025
