Self-Supervised Learning of Pretext-Invariant Representations
Ishan Misra, Laurens van der Maaten
Abstract
The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations. Many pretext tasks lead to representations that are covariant with image transformations. We argue that, instead, semantic representations ought to be invariant under such transformations. Specifically, we develop Pretext-Invariant Representation Learning (PIRL, pronounced as "pearl") that learns invariant representations based on pretext tasks. We use PIRL with a commonly used pretext task that involves solving jigsaw puzzles. We find that PIRL substantially improves the semantic quality of the learned image representations. Our approach sets a new stateof-the-art in self-supervised learning from images on several popular benchmarks for self-supervised learning. Despite being unsupervised, PIRL outperforms supervised pre-training in learning image representations for object detection. Altogether, our results demonstrate the potential of self-supervised representations with good invariance properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92ec9b0c-df3a-4480-8d65-8d45d2bb26ddCited by top-tier papers375
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Builds on7
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 854 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
Related papers
- Equivariant Self-Supervised Learning: Encouraging Equivariance in RepresentationsRumen Dangovski, Li Jing, Charlotte Loh, Seungwook Han et al.ICLR 2022 · 54 citations
- Self-supervised Transformation Learning for Equivariant RepresentationsJaemyung Yu, Jaehyun Choi, Dong-Jae Lee, Hyeong Gwon Hong et al.NeurIPS 2024 · 10 citations
- Solving Mixed-Modal Jigsaw Puzzle for Fine-Grained Sketch-Based Image RetrievalKaiyue Pang, Yongxin Yang, Timothy M. Hospedales, Tao Xiang et al.CVPR 2020
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li et al.CVPR 2022 · 29 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
