Compressed Video Contrastive Learning
Yuqi Huo, Mingyu Ding, Haoyu Lu, Nanyi Fei, Zhiwu Lu, Ji-Rong Wen, Ping Luo
Abstract
This work concerns self-supervised video representation learning (SSVRL), one topic that has received much attention recently. Since videos are storage-intensive and contain a rich source of visual content, models designed for SSVRL are expected to be storage-and computation-efficient, as well as effective. However, most existing methods only focus on one of the two objectives, failing to consider both at the same time. In this work, for the first time, the seemingly contradictory goals are simultaneously achieved by exploiting compressed videos and capturing mutual information between two input streams. Specifically, a novel Motion Vector based Cross Guidance Contrastive learning approach (MVCGC) is proposed. For storage and computation efficiency, we choose to directly decode RGB frames and motion vectors (that resemble low-resolution optical flows) from compressed videos on-the-fly. To enhance the representation ability of the motion vectors, hence the effectiveness of our method, we design a cross guidance contrastive learning algorithm based on multi-instance InfoNCE loss, where motion vectors can take supervision signals from RGB frames and vice versa. Comprehensive experiments on two downstream tasks show that our MVCGC yields new state-of-the-art while being significantly more efficient than its competitors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan et al.ICCV 2023 · 53 citations
- Compressed Video Prompt TuningBing Li, Jiaxin Chen, Xiuguo Bao, Di HuangNeurIPS 2023 · 11 citations
- From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer LearningYang Liu, Qianqian Xu, Peisong Wen, Siran Dai et al.CVPR 2026 · 2 citations
- RGB No More: Minimally-Decoded JPEG Vision TransformersJeongsoo Park, Justin JohnsonCVPR 2023
- Efficient Motion-Aware Video MLLMZijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo et al.CVPR 2025
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
Related papers
- Self-Supervised Learning of Compressed Video RepresentationsYoungjae Yu, Sangho Lee, Gunhee Kim, Yale SongICLR 2021 · 15 citations
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao et al.ICCV 2021 · 37 citations
- Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation LearningManlin Zhang, Jinpeng Wang, Andy J. MaAAAI 2022 · 9 citations
- Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationAli Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi et al.ICCV 2021 · 65 citations
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
