Gate-Shift Networks for Video Action Recognition
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
摘要
Deep 3D CNNs for video action recognition are designed to learn powerful representations in the joint spatiotemporal feature space. In practice however, because of the large number of parameters and computations involved, they may under-perform in the lack of sufficiently large datasets for training them at scale. In this paper we introduce spatial gating in spatial-temporal decomposition of 3D kernels. We implement this concept with Gate-Shift Module (GSM). GSM is lightweight and turns a 2D-CNN into a highly efficient spatio-temporal feature extractor. With GSM plugged in, a 2D-CNN learns to adaptively route features through time and combine them, at almost no additional parameters and computational overhead. We perform an extensive evaluation of the proposed module to study its effectiveness in video action recognition, achieving state-of-the-art results on Something Something-V1 and Diving48 datasets, and obtaining competitive results on EPIC-Kitchens with far less model complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- MGSampler: An Explainable Sampling Strategy for Video Action RecognitionYuan Zhi, Zhan Tong, Limin Wang, Gangshan WuICCV 2021 · 被引用 89 次
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li 等AAAI 2021 · 被引用 87 次
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li 等ICCV 2023 · 被引用 77 次
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri 等ICLR 2021 · 被引用 70 次
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang 等ICLR 2021 · 被引用 69 次
它引用的顶会 Paper8
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 被引用 204 次
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 被引用 170 次
相关 Paper
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Spatiotemporal Joint Filter Decomposition in 3D Convolutional Neural NetworksZichen Miao, Ze Wang, Xiuyuan Cheng, Qiang QiuNeurIPS 2021 · 被引用 12 次
- Learning Comprehensive Motion Representation for Action RecognitionMingyu Wu, Boyuan Jiang, Donghao Luo, Junchi Yan 等AAAI 2021 · 被引用 12 次
- 3D CNNs With Adaptive Temporal Feature ResolutionsMohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi 等CVPR 2021
