Transformers are Multi-State RNNs
Matanel Oren, Michael Hassid, Yarden Nir, Yossi Adi, Roy Schwartz
Abstract
Transformers are considered conceptually different from the previous generation of stateof-the-art NLP models-recurrent neural networks (RNNs). In this work, we demonstrate that decoder-only transformers can in fact be conceptualized as unbounded multistate RNNs-an RNN variant with unlimited hidden state size. We further show that transformers can be converted into bounded multistate RNNs by fixing the size of their hidden state, effectively compressing their keyvalue cache. We introduce a novel, trainingfree compression policy-Token Omission Via Attention (TOVA). 1 Our experiments with four long range tasks and several LLMs show that TOVA outperforms several baseline compression policies. Particularly, our results are nearly on par with the full model, using in some cases only 1 /8 of the original cache size, which translates to 4.8X higher throughput. Our results shed light on the connection between transformers and RNNs, and help mitigate one of LLMs' most painful computational bottlenecks-the size of their key-value cache. 2 * Equal contribuation 1 Literally "good" in Hebrew. 2 https://github.com/schwartz-lab-NLP/TOVA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers75
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- QUEST: Query-Aware Sparsity for Efficient Long-Context LLM InferenceJiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao et al.ICML 2024 · 316 citations
- MoBA: Mixture of Block Attention for Long-Context LLMsEnzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du et al.NeurIPS 2025 · 219 citations
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan et al.ICML 2024 · 106 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache CompressionPayman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai et al.ICML 2025
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 12 citations
- Numerical Pruning for Efficient Autoregressive ModelsXuan Shen, Zhao Song, Yufa Zhou, Bo Chen et al.AAAI 2025 · 5 citations
- OATS: Outlier-Aware Pruning Through Sparse and Low Rank DecompositionStephen Zhang, Vardan PapyanICLR 2025
- DyCoke: Dynamic Compression of Tokens for Fast Video Large Language ModelsKeda Tao, Can Qin, Haoxuan You, Yang Sui et al.CVPR 2025
