TSM: Temporal Shift Module for Efficient Video Understanding
Ji Lin, Chuang Gan, Song Han
Abstract
The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture temporal relationships; 3D CNN based methods can achieve good performance but are computationally intensive, making it expensive to deploy. In this paper, we propose a generic and effective Temporal Shift Module (TSM) that enjoys both high efficiency and high performance. Specifically, it can achieve the performance of 3D CNN but maintain 2D CNN's complexity. TSM shifts part of the channels along the temporal dimension; thus facilitate information exchanged among neighboring frames. It can be inserted into 2D CNNs to achieve temporal modeling at zero computation and zero parameters. We also extended TSM to online setting, which enables real-time low-latency online video recognition and video object detection. TSM is accurate and efficient: it ranks the first place on the Something-Something leaderboard upon publication; on Jetson Nano and Galaxy Note8, it achieves a low latency of 13ms and 35ms for online video recognition. The code is available at: https://github. com/mit-han-lab/temporal-shift-module .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0a595a5-e674-41e9-8e7b-b2e37dcf939fCited by top-tier papers367
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 842 citations
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu et al.CVPR 2022 · 524 citations
- Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals MeasurementXin Liu, Josh Fromm, Shwetak N. Patel, Daniel McDuffNeurIPS 2020 · 436 citations
Related papers
- Gate-Shift Networks for Video Action RecognitionSwathikiran Sudhakaran, Sergio Escalera, Oswald LanzCVPR 2020
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 2 citations
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang et al.AAAI 2020 · 267 citations
