HOPE for a Robust Parameterization of Long-memory State Space Models
Annan Yu, Michael W. Mahoney, N. Benjamin Erichson
摘要
State-space models (SSMs) that utilize linear, time-invariant (LTI) systems are known for their effectiveness in learning long sequences. To achieve state-of-the-art performance, an SSM often needs a specifically designed initialization, and the training of state matrices is on a logarithmic scale with a very small learning rate. To understand these choices from a unified perspective, we view SSMs through the lens of Hankel operator theory. Building upon it, we develop a new parameterization scheme, called HOPE, for LTI systems that utilizes Markov parameters within Hankel operators. Our approach helps improve the initialization and training stability, leading to a more robust parameterization. We efficiently implement these innovations by nonuniformly sampling the transfer functions of LTI systems, and they require fewer parameters compared to canonical SSMs. When benchmarked against HiPPO-initialized models such as S4 and S4D, an SSM parameterized by Hankel operators demonstrates improved performance on Long-Range Arena (LRA) tasks. Moreover, our new parameterization endows the SSM with non-decaying memory within a fixed time window, which is empirically corroborated by a sequential CIFAR-10 task with padded noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential RecommendationMaolin Wang, Sheng Zhang, Ruocheng Guo, Wanyu Wang 等SIGIR 2025 · 被引用 12 次
- Block-Biased Mamba for Long-Range Sequence ProcessingAnnan Yu, N. Benjamin ErichsonNeurIPS 2025 · 被引用 10 次
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and CompressibilityAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang 等ICLR 2026 · 被引用 9 次
- Generalization Error Analysis for Selective State-Space Models Through the Lens of AttentionArya Honarpisheh, Mustafa Bozdag, Octavia I. Camps, Mario SznaierNeurIPS 2025 · 被引用 6 次
- The Curious Case of In-Training Compression of State Space ModelsMakram Chahine, Philipp Nazari, Daniela Rus, T. Konstantin RuschICLR 2026 · 被引用 4 次
它引用的顶会 Paper15
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 被引用 690 次
- Diagonal State Spaces are as Effective as Structured State SpacesAnkit Gupta, Albert Gu, Jonathan BerantNeurIPS 2022 · 被引用 546 次
相关 Paper
- How to Train your HIPPO: State Space Models with Generalized Orthogonal Basis ProjectionsAlbert Gu, Isys Johnson, Aman Timalsina, Atri Rudra 等ICLR 2023 · 被引用 11 次
- Uncovering the Spectral Bias in Diagonal State Space ModelsRuben Solozabal, Velibor Bojkovic, Hilal AlQuabeh, Kentaro Inui 等NeurIPS 2025 · 被引用 3 次
- Robustifying State-space Models for Long Sequences via Approximate DiagonalizationAnnan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney 等ICLR 2024 · 被引用 18 次
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 被引用 78 次
- Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space ModelsFusheng Liu, Qianxiao LiICLR 2025
