Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
Aleksandar Terzic, Nicolas Menet, Michael Hersche, Thomas Hofmann, Abbas Rahimi
摘要
Modern state-space models (SSMs) often utilize transition matrices which enable efficient computation but pose restrictions on the model's expressivity, as measured in terms of the ability to emulate finite-state automata (FSA). While unstructured transition matrices are optimal in terms of expressivity, they come at a prohibitively high compute and memory cost even for moderate state sizes. We propose a structured sparse parametrization of transition matrices in SSMs that enables FSA state tracking with optimal state size and depth, while keeping the computational cost of the recurrence comparable to that of diagonal SSMs. Our method, PD-SSM, parametrizes the transition matrix as the product of a column one-hot matrix () and a complex-valued diagonal matrix (). Consequently, the computational cost of parallel scans scales linearly with the state size. Theoretically, the model is BIBO-stable and can emulate any -state FSA with one layer of dimension and a linear readout of size , significantly improving on all current structured SSM guarantees. Experimentally, the model significantly outperforms a wide collection of modern SSM variants on various FSA state tracking tasks. On multiclass time-series classification, the performance is comparable to that of neural controlled differential equations, a paradigm explicitly built for time-series analysis. Finally, we integrate PD-SSM into a hybrid Transformer-SSM architecture and demonstrate that the model can effectively track the states of a complex FSA in which transitions are encoded as a set of variable-length English sentences. The code is available at https://github.com/IBM/expressive-sparse-state-space-model
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Why Are Linear RNNs More Parallelizable?William Merrill, Hongjian Jiang, Yanhong Li, Anthony Lin 等ICML 2026 · 被引用 5 次
- On the "Induction Bias" in Sequence ModelsMohammadReza Ebrahimi, Michaël Defferrard, Sunny Panchal, Roland MemisevicICML 2026 · 被引用 3 次
- A Spiking Heterogeneous Harmonic Resonate-and-Fire State Space Model for Time SeriesKartikay Agrawal, Vaishnavi N, Abhijeet Vikram, Vedant Sharma 等ICML 2026
- MIMOMamba: From Scalar Duality to Matrix-Valued AttentionYanbo Li, Richard Cornelius Suwandi, Feng Yin, Yiyong SUN 等ICML 2026
- Rational TransductorsMehryar MohriICML 2026
它引用的顶会 Paper25
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
相关 Paper
- On the Expressiveness and Length Generalization of Selective State Space Models on Regular LanguagesAleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann 等AAAI 2025 · 被引用 8 次
- Effectively Modeling Time Series with Simple Discrete State SpacesMichael Zhang, Khaled Kamal Saab, Michael Poli, Tri Dao 等ICLR 2023 · 被引用 14 次
- On Structured State-Space DualityJerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh, Weimin Wu 等ICML 2026
- The Expressive Limits of Diagonal SSMs for State-TrackingMehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath ChandarICLR 2026 · 被引用 11 次
- On the Expressiveness of State Space Models via Temporal LogicsEric Alsmann, Lowejatan Noori, Martin LangeICLR 2026 · 被引用 2 次
