Self-Supervised Spatiotemporal Representation Learning by Exploiting Video Continuity
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai, Juwei Lu, Yang Wang
Abstract
Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the video continuity, to obtain supervision signals for self-supervised representation learning. Specifically, we formulate three novel continuity-related pretext tasks, i.e. continuity justification, discontinuity localization, and missing section approximation, that jointly supervise a shared backbone for video representation learning. This self-supervision approach, termed as Continuity Perception Network (CPNet), solves the three tasks altogether and encourages the backbone network to learn local and long-ranged motion and context representations. It outperforms prior arts on multiple downstream tasks, such as action recognition, video retrieval, and action localization. Additionally, the video continuity can be complementary to other coarse-grained video properties for representation learning, and integrating the proposed pretext task to prior arts can yield much performance gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsHanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang et al.NeurIPS 2024 · 116 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
- Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsTao Zhong, Zhixiang Chi, Li Gu, Yang Wang et al.NeurIPS 2022 · 70 citations
- MetaGCD: Learning to Continually Learn in Generalized Category DiscoveryYanan Wu, Zhixiang Chi, Yang Wang, Songhe FengICCV 2023 · 50 citations
- Test-Time Domain Adaptation by Learning Domain-Aware Batch NormalizationYanan Wu, Zhixiang Chi, Yang Wang, Konstantinos N. Plataniotis et al.AAAI 2024 · 41 citations
Builds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long et al.AAAI 2021 · 140 citations
Related papers
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Self-Supervised Learning of Compressed Video RepresentationsYoungjae Yu, Sangho Lee, Gunhee Kim, Yale SongICLR 2021 · 15 citations
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu et al.ICCV 2021 · 55 citations
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
