State-Free Inference of State-Space Models: The Transfer Function Approach
Rom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith, Ramin M. Hasani, Mathias Lechner, Qi An, Christopher Ré, Hajime Asama, Stefano Ermon, Taiji Suzuki, Michael Poli, Atsushi Yamashita
Abstract
We approach designing a state-space model for deep learning applications through its dual representation, the transfer function, and uncover a highly efficient sequence parallel inference algorithm that is state-free: unlike other proposed algorithms, state-free inference does not incur any significant memory or computational cost with an increase in state size. We achieve this using properties of the proposed frequency domain transfer function parametrization, which enables direct computation of its corresponding convolutional kernel's spectrum via a single Fast Fourier Transform. Our experimental results across multiple sequence lengths and state sizes illustrates, on average, a 35% training speed improvement over S4 layers -- parametrized in time-domain -- on the Long Range Arena benchmark, while delivering state-of-the-art downstream performances over other attention-free approaches. Moreover, we report improved perplexity in language modeling over a long convolutional Hyena baseline, by simply introducing our transfer function parametrization. Our code is available at https://github.com/ruke1ire/RTF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 854a0335-836a-42e1-b03b-1f293fdf7833Cited by top-tier papers11
- Towards Scalable and Stable Parallelization of Nonlinear RNNsXavier Gonzalez, Andrew Warrington, Jimmy T. H. Smith, Scott W. LindermanNeurIPS 2024 · 47 citations
- Parallelization of Non-linear State-Space Models: Scaling Up Liquid-Resistance Liquid-Capacitance Networks for Efficient Sequence ModelingMónika Farsang, Radu GrosuNeurIPS 2025 · 14 citations
- Layer-Adaptive State Pruning for Deep State Space ModelsMinseon Gwak, Seongrok Moon, Joohwan Ko, PooGyeon ParkNeurIPS 2024 · 13 citations
- Adaptive Mamba Neural OperatorsZeyuan Song, Zheyu JiangICLR 2026 · 6 citations
- Uncovering the Spectral Bias in Diagonal State Space ModelsRuben Solozabal, Velibor Bojkovic, Hilal AlQuabeh, Kentaro Inui et al.NeurIPS 2025 · 3 citations
Builds on20
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 850 citations
Related papers
- Laughing Hyena Distillery: Extracting Compact Recurrences From ConvolutionsStefano Massaroli, Michael Poli, Daniel Y. Fu, Hermann Kumbong et al.NeurIPS 2023 · 31 citations
- Simple Hardware-Efficient Long Convolutions for Sequence ModelingDaniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas et al.ICML 2023 · 72 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
