Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization
Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, Weiyao Lin
Abstract
The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics and neglected lower-level representations and their temporal relationship which are crucial for general video understanding. To address these challenges, this paper proposes a multi-level feature optimization framework to improve the generalization and temporal modeling ability of learned video representations. Concretely, high-level features obtained from naive and prototypical contrastive learning are utilized to build distribution graphs, guiding the process of low-level and mid-level feature learning. We also devise a simple temporal modeling module from multi-level features to enhance motion pattern learning. Experiments demonstrate that multi-level feature optimization with the graph constraint and temporal modeling can greatly improve the representation ability in video understanding. Code is available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d4d3b67-daa1-4b34-82da-8c3ebc9ed91aCited by top-tier papers11
- SPAct: Self-supervised Privacy Preservation for Action RecognitionIshan Rajendrakumar Dave, Chen Chen, Mubarak ShahCVPR 2022 · 62 citations
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian et al.CVPR 2022 · 54 citations
- Probabilistic Representations for Video Contrastive LearningJungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon SohnCVPR 2022 · 42 citations
- Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in VideosRui Qian, Shuangrui Ding, Xian Liu, Dahua LinICCV 2023 · 23 citations
- Dual Contrastive Learning for Spatio-temporal RepresentationShuangrui Ding, Rui Qian, Hongkai XiongACM MM 2022 · 20 citations
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
Related papers
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed et al.CVPR 2021
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 6 citations
- MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation LearningBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ACM MM 2022 · 6 citations
