Pushdown Layers: Encoding Recursive Structure in Transformer Language Models
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning
摘要
Recursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism. Consequently, Transformer language models poorly capture long-tail recursive structure and exhibit sample-inefficient syntactic generalization. This work introduces Pushdown Layers, a new self-attention layer that models recursive state via a stack tape that tracks estimated depths of every token in an incremental parse of the observed prefix. Transformer LMs with Pushdown Layers are syntactic language models that autoregressively and synchronously update this stack tape as they predict new tokens, in turn using the stack tape to softly modulate attention over tokens—for instance, learning to “skip” over closed constituents. When trained on a corpus of strings annotated with silver constituency parses, Transformers equipped with Pushdown Layers achieve dramatically better and 3-5x more sample-efficient syntactic generalization, while maintaining similar perplexities. Pushdown Layers are a drop-in replacement for standard self-attention. We illustrate this by finetuning GPT2-medium with Pushdown Layers on an automatically parsed WikiText-103, leading to improvements on several GLUE text classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Stack Attention: Improving the Ability of Transformers to Model Hierarchical PatternsBrian DuSell, David ChiangICLR 2024 · 被引用 15 次
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald 等ACL 2024 · 被引用 15 次
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceHoujun Liu, Shikhar Murty, Christopher Manning, Róbert CsordásICML 2026 · 被引用 3 次
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu 等ACL 2024 · 被引用 1 次
- Do LLMs learn a true syntactic universal?John T. Hale, Milos StanojevicEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper9
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Neural Networks and the Chomsky HierarchyGrégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein 等ICLR 2023 · 被引用 45 次
- Learning Hierarchical Structures with Differentiable Nondeterministic StacksBrian DuSell, David ChiangICLR 2022 · 被引用 19 次
- Characterizing intrinsic compositionality in transformers with Tree ProjectionsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningICLR 2023 · 被引用 13 次
- Transformers Learn Shortcuts to AutomataBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy 等ICLR 2023 · 被引用 11 次
相关 Paper
- Recursive Transformer: Boosting Reasoning Ability with State StackKechi Zhang, Ge Li, Jia Li, Huangzhao Zhang 等NeurIPS 2025 · 被引用 1 次
- GiLT: Augmenting Transformer Language Models with Dependency GraphsTianyu Huang, Yida Zhao, Chuyan Zhou, Kewei TuACL 2026
- Staircase Attention for Recurrent Processing of SequencesDa Ju, Stephen Roller, Sainbayar Sukhbaatar, Jason WestonNeurIPS 2022 · 被引用 20 次
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- A Systematic Study of Compositional Syntactic Transformer Language ModelsYida Zhao, Hao Xve, Xiang Hu, Kewei TuACL 2025 · 被引用 1 次
