Titans: Learning to Memorize at Test Time
Ali Behrouz, Peilin Zhong, Vahab Mirrokni
Abstract
Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We show that this neural memory has the advantage of fast parallelizable training while maintaining a fast inference. From a memory perspective, we argue that attention due to its limited context but accurate dependency modeling performs as a short-term memory, while neural memory due to its ability to memorize the data, acts as a long-term, more persistent, memory. Based on these two modules, we introduce a new family of architectures, called Titans, and present three variants to address how one can effectively incorporate memory into this architecture. Our experimental results on language modeling, common-sense reasoning, genomics, and time series tasks show that Titans are more effective than Transformers and recent modern linear recurrent models. They further can effectively scale to larger than 2M context window size with higher accuracy in needle-in-haystack tasks compared to baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47775893-f4a8-4383-a64e-12e5aafb7346Cited by top-tier papers34
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online OptimizationAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniICLR 2026 · 63 citations
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri et al.ICML 2026 · 57 citations
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon LayersZeyuan Allen-ZhuNeurIPS 2025 · 44 citations
Builds on58
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
Related papers
- TNT: Improving Chunkwise Training for Test-Time MemorizationZeman Li, Ali Behrouz, Yuan Deng, Peilin Zhong et al.ICLR 2026 · 7 citations
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- VideoTitans: Scalable Video Prediction with Integrated Short- and Long-term MemoryYoung-Jae Park, Minseok Seo, Hae-Gon JeonNeurIPS 2025 · 4 citations
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence ModelingXiuying Wei, Anunay Yadav, Razvan Pascanu, Caglar GulcehreNeurIPS 2025 · 3 citations
- MELODI: Exploring Memory Compression for Long ContextsYinpeng Chen, DeLesley Hutchins, Aren Jansen, Andrey Zhmoginov et al.ICLR 2025 · 1 citation
