Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
Jerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger, Antonio Orvieto
摘要
Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this challenge, alternative architectures such as linear attention, State Space Models (SSMs), and Recurrent Neural Networks (RNNs) have been considered as more efficient alternatives. While connections between these approaches exist, such models are commonly developed in isolation and there is a lack of theoretical understanding of the shared principles underpinning these architectures and their subtle differences, greatly influencing performance and scalability. In this paper, we introduce the Dynamical Systems Framework (DSF), which allows a principled investigation of all these architectures in a common representation. Our framework facilitates rigorous comparisons, providing new insights on the distinctive characteristics of each model class. For instance, we compare linear attention and selective SSMs, detailing their differences and conditions under which both are equivalent. We also provide principled comparisons between softmax attention and other model classes, discussing the theoretical conditions under which softmax attention can be approximated. Additionally, we substantiate these new insights with empirical validations and mathematical arguments. This shows the DSF's potential to guide the systematic development of future more efficient and scalable foundation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Block-Biased Mamba for Long-Range Sequence ProcessingAnnan Yu, N. Benjamin ErichsonNeurIPS 2025 · 被引用 10 次
- Exploring the Limitations of Mamba in COPY and CoT ReasoningRuifeng Ren, Zhicong Li, Yong LiuEMNLP 2025 · 被引用 6 次
- A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space ModelsMugunthan Shandirasegaran, Hongkang Li, Songyang Zhang, Meng Wang 等ICLR 2026 · 被引用 3 次
- SAS: Simulated Attention ScoreChuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang 等NeurIPS 2025 · 被引用 3 次
- LaTIM: Measuring Latent Token-to-Token Interactions in Mamba ModelsHugo Pitorro, Marcos Vinícius TrevisoACL 2025 · 被引用 2 次
它引用的顶会 Paper28
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapYuhong Chou, Man Yao, Kexin Wang, Yuqi Pan 等NeurIPS 2024 · 被引用 22 次
- Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersWeigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu MoAAAI 2025 · 被引用 3 次
- On Structured State-Space DualityJerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh, Weimin Wu 等ICML 2026
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz 等ICLR 2021 · 被引用 425 次
- On the Expressiveness and Length Generalization of Selective State Space Models on Regular LanguagesAleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann 等AAAI 2025 · 被引用 8 次
