Dependency-based Mixture Language Models
Zhixian Yang, Xiaojun Wan
Abstract
Various models have been proposed to incorporate knowledge of syntactic structures into neural language models. However, previous works have relied heavily on elaborate components for a specific language model, usually recurrent neural network (RNN), which makes themselves unwieldy in practice to fit into other neural language models, such as Transformer and GPT-2. In this paper, we introduce the Dependency-based Mixture Language Models. In detail, we first train neural language models with a novel dependency modeling objective to learn the probability distribution of future dependent tokens given context. We then formulate the next-token probability by mixing the previous dependency modeling probability distributions with selfattention. Extensive experiments and human evaluations show that our method can be easily and effectively applied to different neural language models while improving neural text generation on various tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feaeb22b-64a1-4e2a-a858-6459d4f342d6Cited by top-tier papers2
- Relation-Constrained Decoding for Text GenerationXiang Chen, Zhixian Yang, Xiaojun WanNeurIPS 2022 · 7 citations
- Causal Graphical Models for Vision-Language Compositional UnderstandingFiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi et al.ICLR 2025
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 46 citations
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachWenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell et al.ACL 2020 · 15 citations
Related papers
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language ModelsYida Zhao, Chao Lou, Kewei TuACL 2024
- Pretraining with Artificial Language: Studying Transferable Knowledge in Language ModelsRyokan Ri, Yoshimasa TsuruokaACL 2022 · 40 citations
- Trading Complexity for Expressivity Through Structured Generalized Linear Token MixingErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICML 2026 · 1 citation
- Recurrent Hierarchical Topic-Guided RNN for Language GenerationDandan Guo, Bo Chen, Ruiying Lu, Mingyuan ZhouICML 2020 · 20 citations
