Lune

NeurIPS2025顶会

DeltaFormer: Unlock the state space of Transformer

Mingyu Xu, Tenglong Ao, Jiaao He, Jianqiao Lu, Guang Shi, Mingwu Zheng

2025年份
3被引次数

摘要

In recent years, large language models built around the Transformer architecture have achieved breakthrough progress in many fields. At the same time, certain weaknesses in these models have prompted further reflection, with the most fundamental concerns centered on the Transformer architecture itself. The Transformer offers high parallelism and can fully exploit the computing power of GPUs, which has enabled it to replace models such as LSTM over the past few years. However, high parallelism is not a free advantage, as it imposes fundamental limits on model performance. In particular, the problems that the logarithmic-precision Transformer architecture can solve are strictly bounded within the class T C 0 . Many important tasks are generally considered outside T C 0 , including Python code execution, entity tracking, chess, and other state-tracking problems. Meanwhile, recent state-space methods based on the Delta Rule have been able to surpass the T C 0 limitations of the Transformer, but these approaches suffer from fixed-size state spaces and perform poorly on many tasks. To address this, we re-examine the Transformer from the perspective of a state space with kernel functions, and propose an improved architecture called DeltaFormer. We theoretically and empirically demonstrate that this new architecture can overcome the inherent T C 0 expressivity limitations of standard Transformers, while remaining at least as effective in language modeling tasks. We hope our work will inspire the design of more expressive models.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper40

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖