ATLAS: Learning to Optimally Memorize the Context at Test Time
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, Vahab Mirrokni
Abstract
Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale. Their quadratic memory and time complexity, however, bound their applicability in longer sequences and so has motivated researchers to explore effective alternative architectures such as modern recurrent neural networks (a.k.a long-term recurrent memory module). Despite their recent success in diverse downstream tasks, they struggle in tasks that requires long context understanding and extrapolation to longer sequences. We observe that these shortcomings come from three disjoint aspects in their design: (1) limited memory capacity that is bounded by the architecture of memory and feature mapping of the input; (2) online nature of update, i.e., optimizing the memory only with respect to the last input; and (3) less expressive management of their fixed-size memory. To enhance all these three aspects, we present Atlas, a long-term memory module with high capacity that learns to memorize the context by optimizing the memory based on the current and past tokens, overcoming the online nature of long-term memory models. Building on this insight, we present a new family of Transformer-like architectures, called DeepTransformers, that are strict generalizations of the original Transformer architecture. Our experimental results on language modeling, common-sense reasoning, recall-intensive, and long-context understanding tasks show that Atlas surpasses the performance of Transformers and recent linear recurrent models. Atlas further improves the long context performance of Titans, achieving +80% accuracy in 10M context length of BABILong benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd9bf5f8-b61e-45f5-a96b-34c4e3692a52Cited by top-tier papers8
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 368 citations
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- TNT: Improving Chunkwise Training for Test-Time MemorizationZeman Li, Ali Behrouz, Yuan Deng, Peilin Zhong et al.ICLR 2026 · 7 citations
- PERK: Long-Context Reasoning as Parameter-Efficient Test-Time LearningZeming Chen, Angelika Romanou, Gail Weiss, Antoine BosselutICLR 2026 · 4 citations
- Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear RecurrencesNeehal Tumma, Noel Loo, Daniela RusICML 2026 · 3 citations
Builds on33
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
Related papers
- Memory Caching: RNNs with Growing MemoryAli Behrouz, Zeman Li, Yuan Deng, Peilin Zhong et al.ICML 2026 · 10 citations
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online OptimizationAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniICLR 2026 · 63 citations
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention ModelsJiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li et al.ICLR 2026 · 6 citations
- Forgetting Transformer: Softmax Attention with a Forget GateZhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron C. CourvilleICLR 2025
