Recurring the Transformer for Video Action Recognition
Jiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang, Jiajun Shen, Dahai Yu
Abstract
Existing video understanding approaches, such as 3D convolutional neural networks and Transformer-Based methods, usually process the videos in a clip-wise manner; hence huge GPU memory is needed and fixed-length video clips are usually required. To alleviate those issues, we introduce a novel Recurrent Vision Transformer (RViT) framework based on spatial-temporal representation learning to achieve the video action recognition task. Specifically, the proposed RViT is equipped with an attention gate to build interaction between current frame input and previous hidden state, thus aggregating the global level interframe features through the hidden state temporally. RViT is executed recurrently to process a video by giving the current frame and previous hidden state. The RViT can capture both spatial and temporal features because of the attention gate and recurrent execution. Besides, the proposed RViT can work on variant-length video clips properly without requiring large GPU memory thanks to the frame by frame processing flow. Our experiment results demonstrate that RViT can achieve state-of-the-art performance on various datasets for the video recognition task. Specifically, RViT can achieve a top-1 accuracy of 81.5% on Kinetics-400, 92.31% on Jester, 67.9% on Something-Something-V2, and an mAP accuracy of 66.1% on Charades.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e058d6be-8e5c-4b6d-af9b-756ca7938b27Cited by top-tier papers21
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut ModulationAhmadreza Jeddi, Marco Ciccone, Babak TaatiICLR 2026 · 54 citations
- GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video SegmentationJiewen Yang, Xinpeng Ding, Ziyang Zheng, Xiaowei Xu et al.ICCV 2023 · 32 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
- Are Vision Transformers More Data Hungry Than Newborn Visual Systems?Lalit Pandey, Samantha M. W. Wood, Justin N. WoodNeurIPS 2023 · 24 citations
- MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity RecognitionZiqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing et al.UbiComp 2023 · 22 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 19 citations
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song et al.ICLR 2022
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 40 citations
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai et al.ICCV 2021 · 224 citations
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson et al.CVPR 2026 · 9 citations
