Eventful Transformers: Leveraging Temporal Redundancy in Vision Transformers
Matthew Dutson, Yin Li, Mohit Gupta
Abstract
Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are often applied repeatedly across frames or temporal chunks. In this work, we exploit temporal redundancy between subsequent inputs to reduce the cost of Transformers for video processing. We describe a method for identifying and re-processing only those tokens that have changed significantly over time. Our proposed family of models, Eventful Transformers, can be converted from existing Transformers (often without any re-training) and give adaptive control over the compute cost at runtime. We evaluate our method on large-scale datasets for video object detection (ImageNet VID) and action recognition (EPIC-Kitchens 100). Our approach leads to significant computational savings (on the order of 2-4x) with only minor reductions in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d9597a2-0b17-4fea-96c2-5b6b081c5b28Cited by top-tier papers6
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting AccelerationChanghun Oh, Seongryong Oh, Jinwoo Hwang, Yoonsung Kim et al.ASPLOS 2026 · 6 citations
- Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang et al.ACL 2024 · 5 citations
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim et al.VLDB 2025 · 2 citations
- Quanta Neural Networks: From Photons to PerceptionVarun Sundar, Tianyi Zhang, Sacha Jungerman, Mohit GuptaICCV 2025 · 1 citation
- ReFrame: Layer Caching for Accelerated Inference in Real-Time RenderingLufei Liu, Tor M. AamodtICML 2025
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani et al.NeurIPS 2024 · 29 citations
- Prune Spatio-temporal Tokens by Semantic-aware Temporal AccumulationShuangrui Ding, Peisen Zhao, Xiaopeng Zhang, Rui Qian et al.ICCV 2023 · 28 citations
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision TransformersBowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang et al.NeurIPS 2021 · 209 citations
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin et al.ICLR 2021 · 20 citations
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
