Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang
Abstract
We propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d25b4ac1-6d48-4bca-85d3-f38eee24b90aCited by top-tier papers45
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionNicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi et al.CVPR 2022 · 264 citations
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 169 citations
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long et al.AAAI 2021 · 140 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
Related papers
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Cross-Architecture Self-supervised Video Representation LearningSheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang et al.CVPR 2022 · 23 citations
- Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationYujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu et al.AAAI 2022 · 20 citations
- Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation LearningYuan Yao, Chang Liu, Dezhao Luo, Yu Zhou et al.CVPR 2020
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen et al.AAAI 2022 · 35 citations
