Deep Analysis of CNN-Based Spatio-Temporal Representations for Action Recognition
Chun-Fu (Richard) Chen, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris, John Cohn, Aude Oliva, Quanfu Fan
摘要
In recent years, a number of approaches based on 2D or 3D convolutional neural networks (CNN) have emerged for video action recognition, achieving state-of-the-art results on several large-scale benchmark datasets. In this paper, we carry out in-depth comparative analysis to better understand the differences between these approaches and the progress made by them. To this end, we develop an unified framework for both 2D-CNN and 3D-CNN action models, which enables us to remove bells and whistles and provides a common ground for fair comparison. We then conduct an effort towards a large-scale analysis involving over 300 action recognition models. Our comprehensive analysis reveals that a) a significant leap is made in efficiency for action recognition, but not in accuracy; b) 2D-CNN and 3D-CNN models behave similarly in terms of spatio-temporal representation abilities and transferability. Our codes are available at
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian 等ICCV 2021 · 被引用 356 次
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu 等CVPR 2022 · 被引用 107 次
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko 等NeurIPS 2021 · 被引用 89 次
- DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionThanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo 等CVPR 2022 · 被引用 70 次
- Dynamic Network Quantization for Efficient Video InferenceXimeng Sun, Rameswar Panda, Chun-Fu (Richard) Chen, Aude Oliva 等ICCV 2021 · 被引用 56 次
它引用的顶会 Paper13
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang 等AAAI 2020 · 被引用 267 次
相关 Paper
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang 等ACM MM 2021 · 被引用 3 次
- Recognizing Actions in Videos From Unseen ViewpointsA. J. Piergiovanni, Michael S. RyooCVPR 2021
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li 等AAAI 2021 · 被引用 87 次
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
