Quantifying Memory Utilization with Effective State-Size
Rom N. Parnichkun, Neehal Tumma, Armin W. Thomas, Alessandro Moro, Qi An, Taiji Suzuki, Atsushi Yamashita, Michael Poli, Stefano Massaroli
Abstract
The need to develop a general framework for architecture analysis is becoming increasingly important, given the expanding design space of sequence models. To this end, we draw insights from classical signal processing and control theory, to develop a quantitative measure of memory utilization: the internal mechanisms through which a model stores past information to produce future outputs. This metric, which we call effective state-size (ESS), is tailored to the fundamental class of systems with input-invariant and input-varying linear operators, encompassing a variety of computational units such as variants of attention, convolutions, and recurrences. Unlike prior work on memory utilization, which either relies on raw operator visualizations (e.g. attention maps), or simply the total memory capacity (i.e. cache size) of a model, our metrics provide highly interpretable and actionable measurements. In particular, we show how ESS can be leveraged to improve initialization strategies, inform novel regularizers and advance the performance-efficiency frontier through model distillation. Furthermore, we demonstrate that the effect of context delimiters (such as end-of-speech tokens) on ESS highlights cross-architectural differences in how large language models utilize their available memory to recall information. Overall, we find that ESS provides valuable insights into the dynamics that dictate memory utilization, enabling the design of more efficient and effective sequence models. * Equal contribution. The order of authorship was determined by a coin flip. † Equal senior contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder ProductsJulien Siems, Timur Carstensen, Arber Zela, Frank Hutter et al.NeurIPS 2025 · 75 citations
- Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear RecurrencesNeehal Tumma, Noel Loo, Daniela RusICML 2026 · 3 citations
Builds on30
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- Retrieval-Aware Distillation for Transformer-SSM HybridsAviv Bick, Eric Xing, Albert GuICML 2026 · 4 citations
- IAM: Efficient Inference through Attention Mapping between Different-scale LLMsYi Zhao, Zuchao Li, Hai ZhaoACL 2025 · 3 citations
- Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural NetworksJerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger et al.NeurIPS 2024 · 38 citations
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani et al.NeurIPS 2023 · 14 citations
- ABC: Attention with Bounded-memory ControlHao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama et al.ACL 2022 · 31 citations
