Structural Guidance for Transformer Language Models
Peng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo
Abstract
Transformer-based language models pretrained on large amounts of text data have proven remarkably successful in learning generic transferable linguistic representations. Here we study whether structural guidance leads to more human-like systematic linguistic generalization in Transformer language models without resorting to pre-training on very large amounts of data. We explore two general ideas. The "Generative Parsing" idea jointly models the incremental parse and word sequence as part of the same sequence modeling task. The "Structural Scaffold" idea guides the language model's representation via additional structure loss that separately predicts the incremental constituency parse. We train the proposed models along with a vanilla Transformer language model baseline on a 14 million-token and a 46 million-token subset of the BLLIP dataset, and evaluate models' syntactic generalization performances on SG Test Suites and sized BLiMP. Experiment results across two benchmarks suggest converging evidence that generative structural supervisions can induce more robust and humanlike linguistic generalization in Transformer language models without the need for data intensive pre-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6ef79f0-b8f4-4040-a3c5-aec89f13385dCited by top-tier papers10
- Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR ParsingJiawei Zhou, Tahira Naseem, Ramón Fernandez Astudillo, Young-Suk Lee et al.EMNLP 2021 · 27 citations
- Stack Attention: Improving the Ability of Transformers to Model Hierarchical PatternsBrian DuSell, David ChiangICLR 2024 · 15 citations
- Incorporating Distributions of Discourse Structure for Long Document Abstractive SummarizationDongqi Liu, Yifan Wang, Vera DembergACL 2023 · 12 citations
- Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language ModelsYiwen Wang, Jennifer Hu, Roger Levy, Peng QianEMNLP 2021 · 3 citations
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu et al.ACL 2024 · 1 citation
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachWenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell et al.ACL 2020 · 15 citations
Related papers
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 6 citations
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen et al.EMNLP 2021 · 131 citations
- Explicit Syntactic Guidance for Neural Text GenerationYafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin et al.ACL 2023 · 4 citations
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita et al.EMNLP 2020 · 2 citations
