Self-Supervised Learning of Compressed Video Representations
Youngjae Yu, Sangho Lee, Gunhee Kim, Yale Song
Abstract
Self-supervised learning of video representations has received great attention. Existing methods typically require frames to be decoded before being processed, which increases compute and storage requirements and ultimately hinders large-scale training. In this work, we propose an efficient self-supervised approach to learn video representations by eliminating the expensive decoding step. We use a three-stream video architecture that encodes I-frames and P-frames of a compressed video. Unlike existing approaches that encode I-frames and P-frames individually, we propose to jointly encode them by establishing bidirectional dynamic connections across streams. To enable self-supervised learning, we propose two pretext tasks that leverage the multimodal nature (RGB, motion vector, residuals) and the internal GOP structure of compressed videos. The first task asks our network to predict zeroth-order motion statistics in a spatio-temporal pyramid; the second task asks correspondence types between I-frames and P-frames after applying temporal transformations. We show that our approach achieves competitive performance on compressed video recognition both in supervised and self-supervised regimes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han et al.ACM MM 2022 · 49 citations
- End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionCongcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong et al.CVPR 2022 · 18 citations
- Compressed Video Contrastive LearningYuqi Huo, Mingyu Ding, Haoyu Lu, Nanyi Fei et al.NeurIPS 2021 · 13 citations
- Compressed Video Prompt TuningBing Li, Jiaxin Chen, Xiuguo Bao, Di HuangNeurIPS 2023 · 11 citations
- Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation LearningManlin Zhang, Jinpeng Wang, Andy J. MaAAAI 2022 · 9 citations
Related papers
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen et al.AAAI 2022 · 35 citations
- Multiview Pseudo-Labeling for Semi-supervised Learning from VideoBo Xiong, Haoqi Fan, Kristen Grauman, Christoph FeichtenhoferICCV 2021 · 54 citations
- A Slow-I-Fast-P Architecture for Compressed Video Action RecognitionJiapeng Li, Ping Wei, Yongchi Zhang, Nanning ZhengACM MM 2020 · 49 citations
