Temporal Pyramid Network for Action Recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, Bolei Zhou
Abstract
Visual tempo characterizes the dynamics and the temporal scale of an action. Modeling such visual tempos of different actions facilitates their recognition. Previous works often capture the visual tempo through sampling raw videos at multiple rates and constructing an input-level frame pyramid, which usually requires a costly multi-branch network to handle. In this work we propose a generic Temporal Pyramid Network (TPN) at the feature-level, which can be flexibly integrated into 2D or 3D backbone networks in a plug-andplay manner. Two essential components of TPN, the source of features and the fusion of features, form a feature hierarchy for the backbone so that it can capture action instances at various tempos. TPN also shows consistent improvements over other challenging baselines on several action recognition datasets. Specifically, when equipped with TPN, the 3D ResNet-50 with dense sampling obtains a 2% gain on the validation set of Kinetics-400. A further analysis also reveals that TPN gains most of its improvements on action classes that have large variances in their visual tempos, validating the effectiveness of TPN. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers76
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai et al.ICCV 2021 · 224 citations
- Evidential Deep Learning for Open Set Action RecognitionWentao Bao, Qi Yu, Yu KongICCV 2021 · 204 citations
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
Builds on6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 204 citations
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
Related papers
- Movement Enhancement toward Multi-Scale Video Feature Representation for Temporal Action DetectionZixuan Zhao, Dongqi Wang, Xu ZhaoICCV 2023 · 15 citations
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid NetworkJialin Gao, Zhixiang Shi, Guanshuo Wang, Jiani Li et al.AAAI 2020 · 78 citations
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri et al.ICLR 2021 · 70 citations
