Structural Guidance for Transformer Language Models
Peng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo
摘要
Transformer-based language models pretrained on large amounts of text data have proven remarkably successful in learning generic transferable linguistic representations. Here we study whether structural guidance leads to more human-like systematic linguistic generalization in Transformer language models without resorting to pre-training on very large amounts of data. We explore two general ideas. The "Generative Parsing" idea jointly models the incremental parse and word sequence as part of the same sequence modeling task. The "Structural Scaffold" idea guides the language model's representation via additional structure loss that separately predicts the incremental constituency parse. We train the proposed models along with a vanilla Transformer language model baseline on a 14 million-token and a 46 million-token subset of the BLLIP dataset, and evaluate models' syntactic generalization performances on SG Test Suites and sized BLiMP. Experiment results across two benchmarks suggest converging evidence that generative structural supervisions can induce more robust and humanlike linguistic generalization in Transformer language models without the need for data intensive pre-training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR ParsingJiawei Zhou, Tahira Naseem, Ramón Fernandez Astudillo, Young-Suk Lee 等EMNLP 2021 · 被引用 27 次
- Stack Attention: Improving the Ability of Transformers to Model Hierarchical PatternsBrian DuSell, David ChiangICLR 2024 · 被引用 15 次
- Incorporating Distributions of Discourse Structure for Long Document Abstractive SummarizationDongqi Liu, Yifan Wang, Vera DembergACL 2023 · 被引用 12 次
- Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language ModelsYiwen Wang, Jennifer Hu, Roger Levy, Peng QianEMNLP 2021 · 被引用 3 次
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu 等ACL 2024 · 被引用 1 次
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachWenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell 等ACL 2020 · 被引用 15 次
相关 Paper
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 被引用 6 次
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen 等EMNLP 2021 · 被引用 131 次
- Explicit Syntactic Guidance for Neural Text GenerationYafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin 等ACL 2023 · 被引用 4 次
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita 等EMNLP 2020 · 被引用 2 次
