Repeat After Me: Transformers are Better than State Space Models at Copying
Samy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran Malach
摘要
Transformers are the dominant architecture for sequence modeling, but there is growing interest in models that use a fixed-size latent state that does not depend on the sequence length, which we refer to as "generalized state space models" (GSSMs). In this paper we show that while GSSMs are promising in terms of inference-time efficiency, they are limited compared to transformer models on tasks that require copying from the input context. We start with a theoretical analysis of the simple task of string copying and prove that a two layer transformer can copy strings of exponential length while GSSMs are fundamentally limited by their fixed-size latent state. Empirically, we find that transformers outperform GSSMs in terms of efficiency and generalization on synthetic tasks that require copying the context. Finally, we evaluate pretrained large language models and find that transformer models dramatically outperform state space models at copying and retrieving information from context. Taken together, these results suggest a fundamental gap between transformers and GSSMs on tasks of practical interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper88
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICLR 2026 · 被引用 139 次
- Theoretical Foundations of Deep Selective State-Space ModelsNicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi 等NeurIPS 2024 · 被引用 97 次
- Gated Slot Attention for Efficient Linear-Time Sequence ModelingYu Zhang, Songlin Yang, Rui-Jie Zhu, Yue Zhang 等NeurIPS 2024 · 被引用 93 次
- Hydra: Bidirectional State Space Models Through Generalized Matrix MixersSukjun Hwang, Aakash Sunil Lahoti, Ratish Puduppully, Tri Dao 等NeurIPS 2024 · 被引用 54 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space ModelsEran Malach, Omid Saremi, Sinead Williamson, Arwen Bradley 等ICLR 2026 · 被引用 3 次
- The Expressivity Limits of TransformersMaxime Meyer, Mario Michelessa, Caroline Chaux, Vincent TanICML 2026
- Expressivity-Efficiency Tradeoffs for Hybrid Sequence ModelsJohn Cooper, Mingchen Ma, Ilias Diakonikolas, Frederic SalaICML 2026
- Scaling up the State Size of RNN LLMs for Long-Context ScenariosKai Liu, Jianfei Gao, Kai ChenACL 2025 · 被引用 1 次
- On the "Induction Bias" in Sequence ModelsMohammadReza Ebrahimi, Michaël Defferrard, Sunny Panchal, Roland MemisevicICML 2026 · 被引用 3 次
