Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
Yida Zhao, Chao Lou, Kewei Tu
Abstract
Syntactic Transformer language models aim to 001 achieve better generalization through simulta-002 neously modeling syntax trees and sentences. 003 While prior work has been focusing on adding 004 constituency-based structures to Transformers, 005 we introduce Dependency Transformer Gram-006 mars (DTGs), a new class of Transformer lan-007 guage model with explicit dependency-based 008 inductive bias. DTGs simulate dependency 009 transition systems with constrained attention 010 patterns by modifying attention masks, incor-011 porate the stack information through relative 012 positional encoding, and augment dependency 013 arc representation with a combination of to-014 ken embeddings and operation embeddings. 015 When trained on a dataset of sentences anno-016 tated with dependency trees, DTGs achieve 017 better generalization while maintaining com-018 parable perplexity with Transformer language 019 model baselines. DTGs also outperform re-020 cent constituency-based models, showing that 021 dependency can better guide Transformer lan-022 guage models. Our code will be publicly avail-023 able upon acceptance. 024 1 Introduction 025 Transformer language models have shown strong 026 performance on language modeling tasks and a 027 broad spectrum of downstream tasks (Radford 028
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4db54354-abb0-4b8c-99ea-d119f26e8845Cited by top-tier papers3
- A Systematic Study of Compositional Syntactic Transformer Language ModelsYida Zhao, Hao Xve, Xiang Hu, Kewei TuACL 2025 · 1 citation
- Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMsXinyu Gao, Shaonan Wang, Nai DingACL 2026
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang et al.AAAI 2026
Builds on5
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Learning Hierarchical Structures with Differentiable Nondeterministic StacksBrian DuSell, David ChiangICLR 2022 · 19 citations
- Unsupervised Dependency Graph NetworkYikang Shen, Shawn Tan, Alessandro Sordoni, Peng Li et al.ACL 2022 · 8 citations
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri et al.ACL 2021
- AMR Parsing with Causal Hierarchical Attention and PointersChao Lou, Kewei TuEMNLP 2023
Related papers
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Dependency-based Mixture Language ModelsZhixian Yang, Xiaojun WanACL 2022 · 3 citations
- On the Ability and Limitations of Transformers to Recognize Formal LanguagesSatwik Bhattamishra, Kabir Ahuja, Navin GoyalEMNLP 2020 · 7 citations
- Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningEMNLP 2023
- Retrofitting Structure-aware Transformer Language Model for End TasksHao Fei, Yafeng Ren, Donghong JiEMNLP 2020 · 56 citations
