On the Expressiveness and Length Generalization of Selective State Space Models on Regular Languages
Aleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann, Abu Sebastian, Abbas Rahimi
摘要
Selective state-space models (SSMs) are an emerging alternative to the Transformer, offering the unique advantage of parallel training and sequential inference. Although these models have shown promising performance on a variety of tasks, their formal expressiveness and length generalization properties remain underexplored. In this work, we provide insight into the workings of selective SSMs by analyzing their expressiveness and length generalization performance on regular language tasks, i.e., finite-state automaton (FSA) emulation. We address certain limitations of modern SSM-based architectures by introducing the Selective Dense State-Space Model (SD-SSM), the first selective SSM that exhibits perfect length generalization on a set of various regular language tasks using a single layer. It utilizes a dictionary of dense transition matrices, a softmax selection mechanism that creates a convex combination of dictionary matrices at each time step, and a readout consisting of layer normalization followed by a linear map. We then proceed to evaluate variants of diagonal selective SSMs by considering their empirical performance on commutative and non-commutative automata. We explain the experimental results with theoretical considerations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan 等NeurIPS 2025 · 被引用 36 次
- Structured Sparse Transition Matrices to Enable State Tracking in State-Space ModelsAleksandar Terzic, Nicolas Menet, Michael Hersche, Thomas Hofmann 等NeurIPS 2025 · 被引用 18 次
- Generalization Error Analysis for Selective State-Space Models Through the Lens of AttentionArya Honarpisheh, Mustafa Bozdag, Octavia I. Camps, Mario SznaierNeurIPS 2025 · 被引用 6 次
- Revisiting Bi-Linear State Transitions in Recurrent Neural NetworksReza Ebrahimi, Roland MemisevicNeurIPS 2025 · 被引用 6 次
- Fixed-Point RNNs: Interpolating from Diagonal to DenseSajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, Antonio OrvietoNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 被引用 690 次
相关 Paper
- On the Expressiveness of State Space Models via Temporal LogicsEric Alsmann, Lowejatan Noori, Martin LangeICLR 2026 · 被引用 2 次
- Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothingPeihao Wang, Ruisi Cai, Yuehao Wang, Jiajun Zhu 等ICLR 2025
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space ModelsEran Malach, Omid Saremi, Sinead Williamson, Arwen Bradley 等ICLR 2026 · 被引用 3 次
- Block-State TransformersJonathan Pilault, Mahan Fathi, Orhan Firat, Chris Pal 等NeurIPS 2023 · 被引用 33 次
- Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall CapacityNingyuan Teresa Huang, Miguel Sarabia, Abhinav Moudgil, Pau Rodríguez 等ICML 2025
