DeltaFormer: Unlock the state space of Transformer
Mingyu Xu, Tenglong Ao, Jiaao He, Jianqiao Lu, Guang Shi, Mingwu Zheng
Abstract
In recent years, large language models built around the Transformer architecture have achieved breakthrough progress in many fields. At the same time, certain weaknesses in these models have prompted further reflection, with the most fundamental concerns centered on the Transformer architecture itself. The Transformer offers high parallelism and can fully exploit the computing power of GPUs, which has enabled it to replace models such as LSTM over the past few years. However, high parallelism is not a free advantage, as it imposes fundamental limits on model performance. In particular, the problems that the logarithmic-precision Transformer architecture can solve are strictly bounded within the class T C 0 . Many important tasks are generally considered outside T C 0 , including Python code execution, entity tracking, chess, and other state-tracking problems. Meanwhile, recent state-space methods based on the Delta Rule have been able to surpass the T C 0 limitations of the Transformer, but these approaches suffer from fixed-size state spaces and perform poorly on many tasks. To address this, we re-examine the Transformer from the perspective of a state space with kernel functions, and propose an improved architecture called DeltaFormer. We theoretically and empirically demonstrate that this new architecture can overcome the inherent T C 0 expressivity limitations of standard Transformers, while remaining at least as effective in language modeling tasks. We hope our work will inspire the design of more expressive models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- On the Expressiveness of State Space Models via Temporal LogicsEric Alsmann, Lowejatan Noori, Martin LangeICLR 2026 · 2 citations
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 157 citations
- Limits of Deep Learning: Sequence Modeling through the Lens of Complexity TheoryNikola Zubic, Federico Soldà, Aurelio L. Sulser, Davide ScaramuzzaICLR 2025
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
- Repeat After Me: Transformers are Better than State Space Models at CopyingSamy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran MalachICML 2024 · 176 citations
