Self-Supervised Video Representation Learning by Context and Motion Decoupling
Lianghua Huang, Yu Liu, Bin Wang, Pan Pan, Yinghui Xu, Rong Jin
摘要
A key challenge in self-supervised video representation learning is how to effectively capture motion information besides context bias. While most existing works implicitly achieve this with video-specific pretext tasks (e.g., predicting clip orders, time arrows, and paces), we develop a method that explicitly decouples motion supervision from context bias through a carefully designed pretext task. Specifically, we take the key frames and motion vectors in compressed videos (e.g., in H.264 format) as the supervision sources for context and motion, respectively, which can be efficiently extracted at over 500 fps on CPU. Then we design two pretext tasks that are jointly optimized: a context matching task where a pairwise contrastive loss is cast between video clip and key frame features; and a motion prediction task where clip features, passed through an encoderdecoder network, are used to estimate motion features in a near future. These two tasks use a shared video backbone and separate MLP heads. Experiments show that our approach improves the quality of the learned video representation over previous works, where we obtain absolute gains of 16.0% and 11.1% in video retrieval recall on UCF101 and HMDB51, respectively. Moreover, we find the motion prediction to be a strong regularization for video networks, where using it as an auxiliary task improves the accuracy of action recognition with a margin of 7.4% ∼ 13.8%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud UnderstandingMohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri 等CVPR 2022 · 被引用 286 次
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 等CVPR 2022 · 被引用 111 次
- Learning from Temporal Gradient for Semi-supervised Action RecognitionJunfei Xiao, Longlong Jing, Lin Zhang, Ju He 等CVPR 2022 · 被引用 77 次
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian 等CVPR 2022 · 被引用 54 次
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan 等ICCV 2023 · 被引用 53 次
它引用的顶会 Paper14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 被引用 534 次
相关 Paper
- Self-Supervised Learning of Compressed Video RepresentationsYoungjae Yu, Sangho Lee, Gunhee Kim, Yale SongICLR 2021 · 被引用 15 次
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long 等AAAI 2021 · 被引用 140 次
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 被引用 2 次
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen 等AAAI 2022 · 被引用 35 次
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu 等ICCV 2021 · 被引用 55 次
