Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
Xiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu, Kewei Tu
Abstract
A syntactic language model (SLM) incrementally generates a sentence with its syntactic tree in a left-to-right manner. We present Generative Pretrained Structured Transformers (GPST), an unsupervised SLM at scale capable of being pre-trained from scratch on raw texts with high parallelism. GPST circumvents the limitations of previous SLMs such as relying on gold trees and sequential training. It consists of two components, a usual SLM supervised by a uni-directional language modeling loss, and an additional composition model, which induces syntactic parse trees and computes constituent representations, supervised by a bi-directional language modeling loss. We propose a representation surrogate to enable joint parallel training of the two models in a hard-EM fashion. We pre-train GPST on OpenWebText, a corpus with 9 billion tokens, and demonstrate the superiority of GPST over GPT-2 with a comparable size in numerous tasks covering both language understanding and language generation. Meanwhile, GPST also significantly outperforms existing unsupervised SLMs on left-to-right grammar induction, while holding a substantial acceleration on training. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- A Systematic Study of Compositional Syntactic Transformer Language ModelsYida Zhao, Hao Xve, Xiang Hu, Kewei TuACL 2025 · 1 citation
- Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMsXinyu Gao, Shaonan Wang, Nai DingACL 2026
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language ModelingXiang Hu, Zhihao Teng, Jun Zhao, Wei Wu et al.ICML 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersAndrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman et al.EMNLP 2020 · 27 citations
- Unsupervised Vision-Language Grammar Induction with Shared Structure ModelingBo Wan, Wenjuan Han, Zilong Zheng, Tinne TuytelaarsICLR 2022 · 19 citations
Related papers
- Structural Guidance for Transformer Language ModelsPeng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez AstudilloACL 2021
- Generative Pre-trained Speech Language Model with Efficient Hierarchical TransformerYongxin Zhu, Dan Su, Liqiang He, Linli Xu et al.ACL 2024
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- GraphGPT: Generative Pre-trained Graph Eulerian TransformerQifang Zhao, Weidong Ren, Tianyu Li, Hong Liu et al.ICML 2025
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri et al.ACL 2021
