Overcoming a Theoretical Limitation of Self-Attention
David Chiang, Peter Cholak
摘要
Although transformers are remarkably effective for many tasks, there are some surprisingly easy-looking regular languages that they struggle with. Hahn shows that for languages where acceptance depends on a single input symbol, a transformer’s classification decisions get closer and closer to random guessing (that is, a cross-entropy of 1) as input strings get longer and longer. We examine this limitation using two languages: PARITY, the language of bit strings with an odd number of 1s, and FIRST, the language of bit strings starting with a 1. We demonstrate three ways of overcoming the limitation implied by Hahn’s lemma. First, we settle an open question by constructing a transformer that recognizes PARITY with perfect accuracy, and similarly for FIRST. Second, we use layer normalization to bring the cross-entropy of both models arbitrarily close to zero. Third, when transformers need to focus on a single position, as for FIRST, we find that they can fail to generalize to longer strings; we offer a simple remedy to this problem that also improves length generalization in machine translation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper61
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 被引用 313 次
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin 等ICLR 2024 · 被引用 189 次
- Scaling Laws of RoPE-based ExtrapolationXiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu 等ICLR 2024 · 被引用 130 次
- The Transient Nature of Emergent In-Context Learning in TransformersAaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz, Erin Grant 等NeurIPS 2023 · 被引用 92 次
- Exposing Attention Glitches with Flip-Flop Language ModelingBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy 等NeurIPS 2023 · 被引用 90 次
它引用的顶会 Paper4
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 被引用 183 次
- On the Ability and Limitations of Transformers to Recognize Formal LanguagesSatwik Bhattamishra, Kabir Ahuja, Navin GoyalEMNLP 2020 · 被引用 7 次
- Effects of Parameter Norm Growth During Transformer Training: Inductive Bias from Gradient DescentWilliam Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz 等EMNLP 2021 · 被引用 1 次
相关 Paper
- How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit BiasRuiquan Huang, Yingbin Liang, Jing YangICML 2025
- A Formal Framework for Understanding Length Generalization in TransformersXinting Huang, Andy Yang, Satwik Bhattamishra, Yash Raj Sarrof 等ICLR 2025
- Why are Sensitive Functions Hard for Transformers?Michael Hahn, Mark RofinACL 2024 · 被引用 3 次
- Language Models Need Inductive Biases to Count InductivelyYingshan Chang, Yonatan BiskICLR 2025
- Quantitative Bounds for Length Generalization in TransformersZachary Izzo, Eshaan Nichani, Jason D. LeeICLR 2026 · 被引用 8 次
