Repeat After Me: Transformers are Better than State Space Models at Copying
Samy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran Malach
Abstract
Transformers are the dominant architecture for sequence modeling, but there is growing interest in models that use a fixed-size latent state that does not depend on the sequence length, which we refer to as "generalized state space models" (GSSMs). In this paper we show that while GSSMs are promising in terms of inference-time efficiency, they are limited compared to transformer models on tasks that require copying from the input context. We start with a theoretical analysis of the simple task of string copying and prove that a two layer transformer can copy strings of exponential length while GSSMs are fundamentally limited by their fixed-size latent state. Empirically, we find that transformers outperform GSSMs in terms of efficiency and generalization on synthetic tasks that require copying the context. Finally, we evaluate pretrained large language models and find that transformer models dramatically outperform state space models at copying and retrieving information from context. Taken together, these results suggest a fundamental gap between transformers and GSSMs on tasks of practical interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b99c54ed-c5d2-4104-9b9a-a03d93fa6bbfCited by top-tier papers88
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger et al.ICLR 2026 · 139 citations
- Theoretical Foundations of Deep Selective State-Space ModelsNicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi et al.NeurIPS 2024 · 97 citations
- Gated Slot Attention for Efficient Linear-Time Sequence ModelingYu Zhang, Songlin Yang, Rui-Jie Zhu, Yue Zhang et al.NeurIPS 2024 · 93 citations
- Hydra: Bidirectional State Space Models Through Generalized Matrix MixersSukjun Hwang, Aakash Sunil Lahoti, Ratish Puduppully, Tri Dao et al.NeurIPS 2024 · 54 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space ModelsEran Malach, Omid Saremi, Sinead Williamson, Arwen Bradley et al.ICLR 2026 · 3 citations
- The Expressivity Limits of TransformersMaxime Meyer, Mario Michelessa, Caroline Chaux, Vincent TanICML 2026
- Expressivity-Efficiency Tradeoffs for Hybrid Sequence ModelsJohn Cooper, Mingchen Ma, Ilias Diakonikolas, Frederic SalaICML 2026
- Scaling up the State Size of RNN LLMs for Long-Context ScenariosKai Liu, Jianfei Gao, Kai ChenACL 2025 · 1 citation
- On the "Induction Bias" in Sequence ModelsMohammadReza Ebrahimi, Michaël Defferrard, Sunny Panchal, Roland MemisevicICML 2026 · 3 citations
