Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomes
Artyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail Burtsev
Abstract
In many sequential tasks, a model needs to remember relevant events from the distant past to make correct predictions. Unfortunately, a straightforward application of gradient based training requires intermediate computations to be stored for every element of a sequence. This requires to store prohibitively large intermediate data if a sequence consists of thousands or even millions elements, and as a result, makes learning of very long-term dependencies infeasible. However, the majority of sequence elements can usually be predicted by taking into account only temporally local information. On the other hand, predictions affected by long-term dependencies are sparse and characterized by high uncertainty given only local information. We propose MemUP, a new training method that allows to learn long-term dependencies without backpropagating gradients through the whole sequence at a time. This method can potentially be applied to any recurrent architecture. LSTM network trained with MemUP performs better or comparable to baselines while requiring to store less intermediate data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84a77c7c-3122-4379-bd04-0e5fafe99a64Cited by top-tier papers6
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningEgor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. PanovICLR 2026 · 43 citations
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 15 citations
- Recurrent Action Transformer with MemoryEgor Cherepanov, Aleksei Staroverov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 14 citations
- Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and EvaluationEgor Cherepanov, Nikita Kachaev, Artem Zholus, Alexey K. Kovalev et al.ICLR 2026 · 4 citations
- Q-RAG: Long Context Multi‑Step Retrieval via Value‑Based Embedder TrainingArtyom Y. Sorokin, Nazar Buzun, Alexander Anokhin, Egor Vedernikov et al.ICLR 2026 · 4 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu et al.ACL 2020 · 168 citations
Related papers
- Non-Local Recurrent Neural Memory for Supervised Sequence ModelingCanmiao Fu, Wenjie Pei, Qiong Cao, Chaopeng Zhang et al.ICCV 2019 · 12 citations
- Long Expressive Memory for Sequence ModelingT. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. MahoneyICLR 2022 · 57 citations
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 51 citations
- Training biologically plausible recurrent neural networks on cognitive tasks with long-term dependenciesWayne Soo, Vishwa Goudar, Xiao-Jing WangNeurIPS 2023 · 16 citations
- Practical Real Time Recurrent Learning with a Sparse ApproximationJacob Menick, Erich Elsen, Utku Evci, Simon Osindero et al.ICLR 2021 · 18 citations
