Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
Jerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger, Antonio Orvieto
Abstract
Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this challenge, alternative architectures such as linear attention, State Space Models (SSMs), and Recurrent Neural Networks (RNNs) have been considered as more efficient alternatives. While connections between these approaches exist, such models are commonly developed in isolation and there is a lack of theoretical understanding of the shared principles underpinning these architectures and their subtle differences, greatly influencing performance and scalability. In this paper, we introduce the Dynamical Systems Framework (DSF), which allows a principled investigation of all these architectures in a common representation. Our framework facilitates rigorous comparisons, providing new insights on the distinctive characteristics of each model class. For instance, we compare linear attention and selective SSMs, detailing their differences and conditions under which both are equivalent. We also provide principled comparisons between softmax attention and other model classes, discussing the theoretical conditions under which softmax attention can be approximated. Additionally, we substantiate these new insights with empirical validations and mathematical arguments. This shows the DSF's potential to guide the systematic development of future more efficient and scalable foundation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c16a80e-9917-4b9c-b0b3-3217a5cceadbCited by top-tier papers12
- Block-Biased Mamba for Long-Range Sequence ProcessingAnnan Yu, N. Benjamin ErichsonNeurIPS 2025 · 10 citations
- Exploring the Limitations of Mamba in COPY and CoT ReasoningRuifeng Ren, Zhicong Li, Yong LiuEMNLP 2025 · 6 citations
- A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space ModelsMugunthan Shandirasegaran, Hongkang Li, Songyang Zhang, Meng Wang et al.ICLR 2026 · 3 citations
- SAS: Simulated Attention ScoreChuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang et al.NeurIPS 2025 · 3 citations
- LaTIM: Measuring Latent Token-to-Token Interactions in Mamba ModelsHugo Pitorro, Marcos Vinícius TrevisoACL 2025 · 2 citations
Builds on28
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapYuhong Chou, Man Yao, Kexin Wang, Yuqi Pan et al.NeurIPS 2024 · 22 citations
- Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersWeigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu MoAAAI 2025 · 3 citations
- On Structured State-Space DualityJerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh, Weimin Wu et al.ICML 2026
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
- On the Expressiveness and Length Generalization of Selective State Space Models on Regular LanguagesAleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann et al.AAAI 2025 · 8 citations
