Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
Brian DuSell, David Chiang
摘要
Attention, specifically scaled dot-product attention, has proven effective for natural language, but it does not have a mechanism for handling hierarchical patterns of arbitrary nesting depth, which limits its ability to recognize certain syntactic structures. To address this shortcoming, we propose stack attention: an attention operator that incorporates stacks, inspired by their theoretical connections to context-free languages (CFLs). We show that stack attention is analogous to standard attention, but with a latent model of syntax that requires no syntactic supervision. We propose two variants: one related to deterministic pushdown automata (PDAs) and one based on nondeterministic PDAs, which allows transformers to recognize arbitrary CFLs. We show that transformers with stack attention are very effective at learning CFLs that standard transformers struggle on, achieving strong results on a CFL with theoretically maximal parsing difficulty. We also show that stack attention is more effective at natural language modeling under a constrained parameter budget, and we include results on machine translation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Compositional Generalization Across Distributional Shifts with Sparse Tree OperationsPaul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky 等NeurIPS 2024 · 被引用 8 次
- The Power of Hard Attention Transformers on Data Sequences: A formal language theoretic perspectivePascal Bergsträßer, Chris Köcher, Anthony Widjaja Lin, Georg ZetzscheNeurIPS 2024 · 被引用 7 次
- Recursive Transformer: Boosting Reasoning Ability with State StackKechi Zhang, Ge Li, Jia Li, Huangzhao Zhang 等NeurIPS 2025 · 被引用 1 次
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu 等ACL 2024 · 被引用 1 次
- Do LLMs learn a true syntactic universal?John T. Hale, Milos StanojevicEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper7
- SG-Net: Syntax-Guided Machine Reading ComprehensionZhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan 等AAAI 2020 · 被引用 192 次
- Learning Hierarchical Structures with Differentiable Nondeterministic StacksBrian DuSell, David ChiangICLR 2022 · 被引用 19 次
- Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningEMNLP 2023
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Self-Attention Networks Can Process Bounded Hierarchical LanguagesShunyu Yao, Binghui Peng, Christos H. Papadimitriou, Karthik NarasimhanACL 2021
相关 Paper
- Learning Bounded Context-Free-Grammar via LSTM and the Transformer: Difference and the ExplanationsHui Shi, Sicun Gao, Yuandong Tian, Xinyun Chen 等AAAI 2022 · 被引用 24 次
- Context-free Recognition with TransformersSelim Jerad, Anej Svete, Sophie Hao, Ryan Cotterell 等ICML 2026 · 被引用 3 次
- The Surprising Computational Power of Nondeterministic Stack RNNsBrian DuSell, David ChiangICLR 2023
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Push, Pop, Parallelize: Stack-Augmented Linear Attention via the Delta RuleAnh T Nguyen, Saleh Momeni, Ashutosh Chaubey, Changnan Xiao 等ICML 2026
