Circuit Complexity Bounds for RoPE-based Transformer Architecture
Bo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, Jiahao Zhang
摘要
Characterizing the express power of the Transformer architecture is critical to understanding its capacity limits and scaling law. Recent works provide the circuit complexity bounds to Transformer-like architecture. On the other hand, Rotary Position Embedding () has emerged as a crucial technique in modern large language models, offering superior performance in capturing positional information compared to traditional position embeddings, which shows great potential in application prospects, particularly for the long context scenario. Empirical evidence also suggests that -based Transformer architectures demonstrate greater generalization capabilities compared to conventional Transformer models. In this work, we establish a circuit complexity bound for Transformers with attention. Our key contribution is that we show that unless , a -based Transformer with -precision, layers, hidden dimension cannot solve the Arithmetic formula evaluation problem or the Boolean formula value problem. This result significantly demonstrates the fundamental limitation of the expressivity of the -based Transformer architecture, although it achieves giant empirical success. Our theoretical result not only establishes the complexity bound but also may instruct further work on the -based Transformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LazyDiT: Lazy Learning for the Acceleration of Diffusion TransformersXuan Shen, Zhao Song, Yufa Zhou, Bo Chen 等AAAI 2025 · 被引用 40 次
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan 等NeurIPS 2025 · 被引用 36 次
- Numerical Pruning for Efficient Autoregressive ModelsXuan Shen, Zhao Song, Yufa Zhou, Bo Chen 等AAAI 2025 · 被引用 5 次
- Poly-attention: a general scheme for higher-order self-attentionSayak Chakrabarti, Toniann Pitassi, Josh AlmanICLR 2026 · 被引用 3 次
- Two (narrow) heads are better than (an arbitrarily wide) oneAmanuel Tesfaye, Zeno Kujawa, Rajmohan Rajaraman, Ravi SundaramICLR 2026
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye 等NeurIPS 2023 · 被引用 470 次
相关 Paper
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang 等NeurIPS 2024 · 被引用 56 次
- Beyond Position: the emergence of wavelet-like properties in TransformersValeria Ruscio, Umberto Nanni, Fabrizio SilvestriACL 2025 · 被引用 1 次
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin 等NeurIPS 2025 · 被引用 51 次
- HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationYuhan Chen, Ang Lv, Jian Luan, Bin Wang 等ACL 2025
- Deconstructing Positional Information: From Attention Logits to Training BiasesZihan Gu, Ruoyu Chen, Han Zhang, Hua Zhang 等ICLR 2026 · 被引用 4 次
