Long-term Leap Attention, Short-term Periodic Shift for Video Classification
Hao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo
Abstract
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes times longer sequence than the latter under the current attention of quadratic complexity ( 2 2 ). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy.
However, videos naturally contain redundant information between neighboring frames; thereby, we could potentially suppress attention on visually similar frames in a dilated manner. Based on this hypothesis, we propose the LAPS, a long-term "Leap A ention" (LA), short-term "Periodic Shi " (P-Shift) module for video transformers, with (2 2 ) complexity. Specifically, the "LA" groups long-term frames into pairs, then refactors each discrete pair via attention. The "P-Shift" exchanges features between temporal neighbors to confront the loss of shortterm dynamics. By replacing a vanilla 2D attention with the LAPS, we could adapt a static transformer into a video one, with zero extra parameters and neglectable computation overhead (∼2.6%). Experiments on the standard Kinetics-400 benchmark demonstrate that our LAPS transformer could achieve competitive performances in terms of accuracy, FLOPs, and Params among CNN and transformer SOTAs. We open-source our project in https://github.com/VideoNetworks/LAPS-transformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Redundancy-aware Transformer for Video Question AnsweringYicong Li, Xun Yang, An Zhang, Chun Feng et al.ACM MM 2023 · 23 citations
- 3D Human Pose Estimation with Spatio-Temporal Criss-Cross AttentionZhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong et al.CVPR 2023
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
Related papers
- Space-time Mixing Attention for Video TransformerAdrian Bulat, Juan-Manuel Pérez-Rúa, Swathikiran Sudhakaran, Brais Martínez et al.NeurIPS 2021 · 158 citations
- CMC: Video Transformer Acceleration via CODEC Assisted Matrix CondensingZhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing et al.ASPLOS 2024 · 8 citations
- Eventful Transformers: Leveraging Temporal Redundancy in Vision TransformersMatthew Dutson, Yin Li, Mohit GuptaICCV 2023 · 18 citations
- Video Frame Interpolation with Flow TransformerPan Gao, Haoyue Tian, Jie QinACM MM 2023 · 4 citations
- Token Shift Transformer for Video ClassificationHao Zhang, Yanbin Hao, Chong-Wah NgoACM MM 2021 · 1 citation
