Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
Jinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, Xing Sun
Abstract
One significant factor we expect the video representation learning to capture, especially in contrast with the image representation learning, is the object motion. However, we found that in the current mainstream video datasets, some action categories are highly related with the scene where the action happens, making the model tend to degrade to a solution where only the scene information is encoded. For example, a trained model may predict a video as playing football simply because it sees the field, neglecting that the subject is dancing as a cheerleader on the field. This is against our original intention towards the video representation learning and may bring scene bias on a different dataset that can not be ignored. In order to tackle this problem, we propose to decouple the scene and the motion (DSM) with two simple operations, so that the model attention towards the motion information is better paid. Specifically, we construct a positive clip and a negative clip for each video. Compared to the original video, the positive/negative is motion-untouched/broken but scene-broken/untouched by Spatial Local Disturbance and Temporal Local Disturbance. Our objective is to pull the positive closer while pushing the negative farther to the original clip in the latent space. In this way, the impact of the scene is weakened while the temporal sensitivity of the network is further enhanced. We conduct experiments on two tasks with various backbones and different pre-training datasets, and find that our method surpass the SOTA methods with a remarkable 8.1% and 8.8% improvement towards action recognition task on the UCF101 and HMDB51 datasets respectively using the same backbone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian et al.CVPR 2022 · 54 citations
- Learning to Refactor Action and Co-occurrence Features for Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Nanning Zheng et al.CVPR 2022 · 43 citations
- Enhancing Self-supervised Video Representation Learning via Multi-level Feature OptimizationRui Qian, Yuxi Li, Huabin Liu, John See et al.ICCV 2021 · 43 citations
- Probabilistic Representations for Video Contrastive LearningJungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon SohnCVPR 2022 · 42 citations
- Cross-Architecture Self-supervised Video Representation LearningSheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang et al.CVPR 2022 · 23 citations
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang et al.CVPR 2021
Related papers
- Dual Contrastive Learning for Spatio-temporal RepresentationShuangrui Ding, Rui Qian, Hongkai XiongACM MM 2022 · 20 citations
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation LearningJinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin et al.CVPR 2021
- Masked Motion Encoding for Self-Supervised Video Representation LearningXinyu Sun, Peihao Chen, Liangwei Chen, Changhao Li et al.CVPR 2023
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
