Self-Supervised Learning of Video-Induced Visual Invariances
Michael Tschannen, Josip Djolonga, Marvin Ritter, Aravindh Mahendran, Neil Houlsby, Sylvain Gelly, Mario Lucic
Abstract
We propose a general framework for self-supervised learning of transferable visual representations based on Video-Induced Visual Invariances (VIVI). We consider the implicit hierarchy present in the videos and make use of (i) frame-level invariances (e.g. stability to color and contrast perturbations), (ii) shot/clip-level invariances (e.g. robustness to changes in object orientation and lighting conditions), and (iii) video-level invariances (semantic relationships of scenes across shots/clips), to define a holistic selfsupervised loss. Training models using different variants of the proposed framework on videos from the YouTube-8M (YT8M) data set, we obtain state-of-the-art self-supervised transfer learning results on the 19 diverse downstream tasks of the Visual Task Adaptation Benchmark (VTAB), using only 1000 labels per task. We then show how to co-train our models jointly with labeled images, outperforming an ImageNet-pretrained ResNet-50 by 0.8 points with 10× fewer labeled images, as well as the previous best supervised model by 3.7 points using the full ImageNet data set. METHOD MEAN NAT. SPEC. STR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Contrastive learning of global and local features for medical image segmentation with limited annotationsKrishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender KonukogluNeurIPS 2020 · 714 citations
- Big Self-Supervised Models Advance Medical Image ClassificationShekoofeh Azizi, Basil Mustafa, Fiona Ryan, Zachary Beaver et al.ICCV 2021 · 695 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
Builds on6
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
Related papers
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li et al.CVPR 2022 · 29 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Self-Supervised Representation Learning from Flow EquivarianceYuwen Xiong, Mengye Ren, Wenyuan Zeng, Raquel Urtasun WaabiICCV 2021 · 32 citations
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 35 citations
- What Should Not Be Contrastive in Contrastive LearningTete Xiao, Xiaolong Wang, Alexei A. Efros, Trevor DarrellICLR 2021 · 338 citations
