No Frame Left Behind: Full Video Action Recognition
Xin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij, Jan C. van Gemert
Abstract
Not all video frames are equally informative for recognizing an action. It is computationally infeasible to train deep networks on all video frames when actions develop over hundreds of frames. A common heuristic is uniformly sampling a small number of video frames and using these to recognize the action. Instead, here we propose full video action recognition and consider all video frames. To make this computational tractable, we first cluster all frame activations along the temporal dimension based on their similarity with respect to the classification task, and then temporally aggregate the frames in the clusters into a smaller number of representations. Our method is end-to-end trainable and computationally efficient as it relies on temporally localized clustering in combination with fast Hamming distances in feature space. We evaluate on UCF101, HMDB51, Breakfast, and Something-Something V1 and V2, where we compare favorably to existing heuristic frame sampling methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Revisiting the "Video" in Video-Language UnderstandingShyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu et al.CVPR 2022 · 121 citations
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu et al.CVPR 2022 · 107 citations
- DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionThanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo et al.CVPR 2022 · 70 citations
- Cross-Modal Learning with 3D Deformable Attention for Action RecognitionSangwon Kim, Dasom Ahn, ByoungChul KoICCV 2023 · 49 citations
- How Do You Do It? Fine-Grained Action Understanding with Pseudo-AdverbsHazel Doughty, Cees G. M. SnoekCVPR 2022 · 16 citations
Builds on8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
Related papers
- FASTER Recurrent Networks for Efficient Video ClassificationLinchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang et al.AAAI 2020 · 60 citations
- SMART Frame Selection for Action RecognitionShreyank N. Gowda, Marcus Rohrbach, Laura Sevilla-LaraAAAI 2021 · 171 citations
- Temporally-Weighted Hierarchical Clustering for Unsupervised Action SegmentationM. Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba et al.CVPR 2021
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang et al.CVPR 2020
- 3D CNNs With Adaptive Temporal Feature ResolutionsMohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi et al.CVPR 2021
