Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation Learning
Jinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin, Andy J. Ma, Hao Cheng, Pai Peng, Feiyue Huang, Rongrong Ji, Xing Sun
Abstract
Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some of the current methods tend to cheat from the background, i.e., the prediction is highly dependent on the video background instead of the motion, making the model vulnerable to background changes. To mitigate the model reliance towards the background, we propose to remove the background impact by adding the background. That is, given a video, we randomly select a static frame and add it to every other frames to construct a distracting video sample. Then we force the model to pull the feature of the distracting video and the feature of the original video closer, so that the model is explicitly restricted to resist the background influence, focusing more on the motion changes. We term our method as Background Erasing (BE). It is worth noting that the implementation of our method is so simple and neat and can be added to most of the SOTA methods without much efforts. Specifically, BE brings 16.4% and 19.1% improvements with MoCo on the severely biased datasets UCF101 and HMDB51, and 14.5% improvement on the less biased dataset Diving48.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan et al.CVPR 2022 · 111 citations
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko et al.NeurIPS 2021 · 89 citations
- Unsupervised Pre-training for Temporal Action Localization TasksCan Zhang, Tianyu Yang, Junwu Weng, Meng Cao et al.CVPR 2022 · 56 citations
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu et al.ICCV 2021 · 55 citations
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian et al.CVPR 2022 · 54 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
Related papers
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionJinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu et al.AAAI 2021 · 70 citations
- Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation LearningManlin Zhang, Jinpeng Wang, Andy J. MaAAAI 2022 · 9 citations
- VideoMoCo: Contrastive Video Representation Learning With Temporally Adversarial ExamplesTian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang et al.CVPR 2021
- Dual Contrastive Learning for Spatio-temporal RepresentationShuangrui Ding, Rui Qian, Hongkai XiongACM MM 2022 · 20 citations
