Temporal Latent Bottleneck: Synthesis of Fast and Slow Processing Mechanisms in Sequence Learning
Aniket Didolkar, Kshitij Gupta, Anirudh Goyal, Nitesh B. Gundavarapu, Alex Lamb, Nan Rosemary Ke, Yoshua Bengio
Abstract
Recurrent neural networks have a strong inductive bias towards learning temporally compressed representations, as the entire history of a sequence is represented by a single vector. By contrast, Transformers have little inductive bias towards learning temporally compressed representations, as they allow for attention over all previously computed elements in a sequence. Having a more compressed representation of a sequence may be beneficial for generalization, as a high-level representation may be more easily re-used and re-purposed and will contain fewer irrelevant details. At the same time, excessive compression of representations comes at the cost of expressiveness. We propose a solution which divides computation into two streams. A slow stream that is recurrent in nature aims to learn a specialized and compressed representation, by forcing chunks of time steps into a single representation which is divided into multiple vectors. At the same time, a fast stream is parameterized as a Transformer to process chunks consisting of time-steps conditioned on the information in the slow-stream. In the proposed approach we hope to gain the expressiveness of the Transformer, while encouraging better compression and structuring of representations in the slow stream. We show the benefits of the proposed method in terms of improved sample efficiency and generalization performance as compared to various competitive baselines for visual perception and sequential decision making tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c5225ee-d9ae-4386-bfde-fd2fef9f4532Cited by top-tier papers8
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- Generalizable, real-time neural decoding with hybrid state-space modelsAvery Hee-Woon Ryoo, Nanda H. Krishna, Ximeng Mao, Mehdi Azabou et al.NeurIPS 2025 · 16 citations
- SutraNets: Sub-series Autoregressive Networks for Long-Sequence, Probabilistic ForecastingShane Bergsma, Timothy Zeyl, Lei GuoNeurIPS 2023 · 14 citations
- iRadar: Synthesizing Millimeter-Waves from Wearable Inertial Inputs for Human Gesture SensingHuanqi Yang, Mingda Han, Xinyue Li, Di Duan et al.INFOCOM 2025 · 9 citations
- Recursion in Recursion: Two-Level Nested Recursion for Length Generalization with ScalabilityJishnu Ray Chowdhury, Cornelia CarageaNeurIPS 2023 · 7 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal et al.NeurIPS 2025 · 9 citations
- TRACE: A Fast Transformer-based General-Purpose Lossless CompressorYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueWWW 2022 · 60 citations
- Accelerating General-purpose Lossless Compression via Simple and Scalable ParameterizationYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueACM MM 2022 · 8 citations
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical PerspectiveAlireza Mousavi-Hosseini, Clayton Sanford, Denny Wu, Murat A. ErdogduNeurIPS 2025 · 6 citations
- Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li et al.AAAI 2025 · 4 citations
