Eventful Transformers: Leveraging Temporal Redundancy in Vision Transformers
Matthew Dutson, Yin Li, Mohit Gupta
摘要
Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are often applied repeatedly across frames or temporal chunks. In this work, we exploit temporal redundancy between subsequent inputs to reduce the cost of Transformers for video processing. We describe a method for identifying and re-processing only those tokens that have changed significantly over time. Our proposed family of models, Eventful Transformers, can be converted from existing Transformers (often without any re-training) and give adaptive control over the compute cost at runtime. We evaluate our method on large-scale datasets for video object detection (ImageNet VID) and action recognition (EPIC-Kitchens 100). Our approach leads to significant computational savings (on the order of 2-4x) with only minor reductions in accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting AccelerationChanghun Oh, Seongryong Oh, Jinwoo Hwang, Yoonsung Kim 等ASPLOS 2026 · 被引用 6 次
- Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang 等ACL 2024 · 被引用 5 次
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim 等VLDB 2025 · 被引用 2 次
- Quanta Neural Networks: From Photons to PerceptionVarun Sundar, Tianyi Zhang, Sacha Jungerman, Mohit GuptaICCV 2025 · 被引用 1 次
- ReFrame: Layer Caching for Accelerated Inference in Real-Time RenderingLufei Liu, Tor M. AamodtICML 2025
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Mixture of Nested Experts: Adaptive Processing of Visual TokensGagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani 等NeurIPS 2024 · 被引用 29 次
- Prune Spatio-temporal Tokens by Semantic-aware Temporal AccumulationShuangrui Ding, Peisen Zhao, Xiaopeng Zhang, Rui Qian 等ICCV 2023 · 被引用 28 次
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision TransformersBowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang 等NeurIPS 2021 · 被引用 209 次
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin 等ICLR 2021 · 被引用 20 次
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu 等ICCV 2023 · 被引用 31 次
