Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video Representation
Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu
Abstract
Spatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e544cc8c-4187-439f-8ee7-6cceb879df57Cited by top-tier papers4
- Neural Temporal Walks: Motif-Aware Representation Learning on Continuous-Time Dynamic GraphsMing Jin, Yuan-Fang Li, Shirui PanNeurIPS 2022 · 130 citations
- Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation LearningMinghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu et al.ACM MM 2023 · 6 citations
- Spatio-Temporal Crop Aggregation for Video Representation LearningSepehr Sameni, Simon Jenni, Paolo FavaroICCV 2023 · 4 citations
- Fine-grained Key-Value Memory Enhanced Predictor for Video Representation LearningXiaojie Li, Jianlong Wu, Shaowei He, Shuo Kang et al.ACM MM 2023 · 1 citation
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
Related papers
- Contextualized Spatio-Temporal Contrastive Learning with Self-SupervisionLiangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong et al.CVPR 2022 · 24 citations
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation LearningYuan Yao, Chang Liu, Dezhao Luo, Yu Zhou et al.CVPR 2020
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
