Grouped Spatial-Temporal Aggregation for Efficient Action Recognition
Chenxu Luo, Alan L. Yuille
Abstract
Temporal reasoning is an important aspect of video analysis. 3D CNN shows good performance by exploring spatial-temporal features jointly in an unconstrained way, but it also increases the computational cost a lot. Previous works try to reduce the complexity by decoupling the spatial and temporal filters. In this paper, we propose a novel decomposition method that decomposes the feature channels into spatial and temporal groups in parallel. This decomposition can make two groups focus on static and dynamic cues separately. We call this grouped spatial-temporal aggregation (GST). This decomposition is more parameter-efficient and enables us to quantitatively analyze the contributions of spatial and temporal features in different layers. We verify our model on several action recognition tasks that require temporal reasoning and show its effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian et al.ICCV 2021 · 356 citations
- ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningJunting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao et al.NeurIPS 2022 · 290 citations
- MGSampler: An Explainable Sampling Strategy for Video Action RecognitionYuan Zhi, Zhan Tong, Limin Wang, Gangshan WuICCV 2021 · 89 citations
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li et al.AAAI 2021 · 87 citations
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
Builds on1
Related papers
- Spatiotemporal Joint Filter Decomposition in 3D Convolutional Neural NetworksZichen Miao, Ze Wang, Xiuyuan Cheng, Qiang QiuNeurIPS 2021 · 12 citations
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- Gate-Shift Networks for Video Action RecognitionSwathikiran Sudhakaran, Sergio Escalera, Oswald LanzCVPR 2020
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang et al.ACM MM 2021 · 3 citations
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang et al.ACM MM 2020 · 1 citation
