A Slow-I-Fast-P Architecture for Compressed Video Action Recognition
Jiapeng Li, Ping Wei, Yongchi Zhang, Nanning Zheng
Abstract
Compressed video action recognition has drawn growing attention for the storage and processing advantages of compressed videos over original raw videos. While the past few years have witnessed remarkable progress in this problem, most existing approaches rely on RGB frames from raw videos and require multi-step training. In this paper, we propose a novel Slow-I-Fast-P (SIFP) neural network model for compressed video action recognition. It consists of the slow I pathway receiving a sparse sampling I-frame clip and the fast P pathway receiving a dense sampling pseudo optical flow clip. An unsupervised estimation method and a new loss function are designed to generate pseudo optical flows in compressed videos. Our model eliminates the dependence on the traditional optical flows calculated from raw videos. The model is trained in an end-to-end way. The proposed method is evaluated on the challenging HMDB51 and UCF101 datasets. The extensive comparison results and ablation studies demonstrate the effectiveness and strength of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8e1da24d-bc00-4477-bb13-7fec6394a0bcCited by top-tier papers8
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan et al.ICCV 2023 · 53 citations
- Self-supervising Action Recognition by Statistical Moment and Subspace DescriptorsLei Wang, Piotr KoniuszACM MM 2021 · 50 citations
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han et al.ACM MM 2022 · 49 citations
- Compressed Video Prompt TuningBing Li, Jiaxin Chen, Xiuguo Bao, Di HuangNeurIPS 2023 · 11 citations
Related papers
- Two-Stream Action Recognition-Oriented Video Super-ResolutionHaochen Zhang, Dong Liu, Zhiwei XiongICCV 2019 · 58 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Self-Supervised Learning of Compressed Video RepresentationsYoungjae Yu, Sangho Lee, Gunhee Kim, Yale SongICLR 2021 · 15 citations
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij et al.CVPR 2021
- Motion-Augmented Self-Training for Video Recognition at Smaller ScaleKirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. SnoekICCV 2021 · 25 citations
