Continual Transformers: Redundancy-Free Attention for Online Inference
Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis
Abstract
Transformers in their common form are inherently limited to operate on whole token sequences rather than on one token at a time. Consequently, their use during online inference on time-series data entails considerable redundancy due to the overlap in successive token sequences. In this work, we propose novel formulations of the Scaled Dot-Product Attention, which enable Transformers to perform efficient online token-by-token inference on a continual input stream. Importantly, our modifications are purely to the order of computations, while the outputs and learned weights are identical to those of the original Transformer Encoder. We validate our Continual Transformer Encoder with experiments on the THUMOS14, TVSeries and GTZAN datasets with remarkable results: Our Continual one- and two-block architectures reduce the floating point operations per prediction by up to 63x and 2.6x, respectively, while retaining predictive performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5d68029-fed1-4953-a8b4-1209109daa2aBuilds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Learning to Iteratively Solve Routing Problems with Dual-Aspect Collaborative TransformerYining Ma, Jingwen Li, Zhiguang Cao, Wen Song et al.NeurIPS 2021 · 230 citations
Related papers
- Sparse Binary Transformers for Multivariate Time Series ModelingMatt Gorbett, Hossein Shirazi, Indrakshi RayKDD 2023 · 18 citations
- VQ-TR: Vector Quantized Attention for Time Series ForecastingKashif Rasul, Andrew Bennett, Pablo Vicente, Umang Gupta et al.ICLR 2024 · 7 citations
- Eventful Transformers: Leveraging Temporal Redundancy in Vision TransformersMatthew Dutson, Yin Li, Mohit GuptaICCV 2023 · 18 citations
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceHoujun Liu, Shikhar Murty, Christopher Manning, Róbert CsordásICML 2026 · 3 citations
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li et al.ACL 2023 · 29 citations
