Temporal Latent Bottleneck: Synthesis of Fast and Slow Processing Mechanisms in Sequence Learning
Aniket Didolkar, Kshitij Gupta, Anirudh Goyal, Nitesh B. Gundavarapu, Alex Lamb, Nan Rosemary Ke, Yoshua Bengio
摘要
Recurrent neural networks have a strong inductive bias towards learning temporally compressed representations, as the entire history of a sequence is represented by a single vector. By contrast, Transformers have little inductive bias towards learning temporally compressed representations, as they allow for attention over all previously computed elements in a sequence. Having a more compressed representation of a sequence may be beneficial for generalization, as a high-level representation may be more easily re-used and re-purposed and will contain fewer irrelevant details. At the same time, excessive compression of representations comes at the cost of expressiveness. We propose a solution which divides computation into two streams. A slow stream that is recurrent in nature aims to learn a specialized and compressed representation, by forcing chunks of time steps into a single representation which is divided into multiple vectors. At the same time, a fast stream is parameterized as a Transformer to process chunks consisting of time-steps conditioned on the information in the slow-stream. In the proposed approach we hope to gain the expressiveness of the Transformer, while encouraging better compression and structuring of representations in the slow stream. We show the benefits of the proposed method in terms of improved sample efficiency and generalization performance as compared to various competitive baselines for visual perception and sequential decision making tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan 等NeurIPS 2023 · 被引用 197 次
- Generalizable, real-time neural decoding with hybrid state-space modelsAvery Hee-Woon Ryoo, Nanda H. Krishna, Ximeng Mao, Mehdi Azabou 等NeurIPS 2025 · 被引用 16 次
- SutraNets: Sub-series Autoregressive Networks for Long-Sequence, Probabilistic ForecastingShane Bergsma, Timothy Zeyl, Lei GuoNeurIPS 2023 · 被引用 14 次
- iRadar: Synthesizing Millimeter-Waves from Wearable Inertial Inputs for Human Gesture SensingHuanqi Yang, Mingda Han, Xinyue Li, Di Duan 等INFOCOM 2025 · 被引用 9 次
- Recursion in Recursion: Two-Level Nested Recursion for Length Generalization with ScalabilityJishnu Ray Chowdhury, Cornelia CarageaNeurIPS 2023 · 被引用 7 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal 等NeurIPS 2025 · 被引用 9 次
- TRACE: A Fast Transformer-based General-Purpose Lossless CompressorYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueWWW 2022 · 被引用 60 次
- Accelerating General-purpose Lossless Compression via Simple and Scalable ParameterizationYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueACM MM 2022 · 被引用 8 次
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical PerspectiveAlireza Mousavi-Hosseini, Clayton Sanford, Denny Wu, Murat A. ErdogduNeurIPS 2025 · 被引用 6 次
- Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li 等AAAI 2025 · 被引用 4 次
