UniVIP: A Unified Framework for Self-Supervised Visual Pre-training
Zhaowen Li, Yousong Zhu, Fan Yang, Wei Li, Chaoyang Zhao, Yingying Chen, Zhiyang Chen, Jiahao Xie, Liwei Wu, Rui Zhao, Ming Tang, Jinqiao Wang
Abstract
Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1365b21b-438a-482d-9084-2aa284cbb62dCited by top-tier papers6
- A Shapelet-based Framework for Unsupervised Multivariate Time Series Representation LearningZhiyu Liang, Jianfeng Zhang, Chen Liang, Hongzhi Wang et al.VLDB 2024 · 19 citations
- Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual TasksZhiyang Chen, Yousong Zhu, Zhaowen Li, Fan Yang et al.NeurIPS 2022 · 17 citations
- Semantics-Consistent Feature Search for Self-Supervised Visual Representation LearningKaiyou Song, Shan Zhang, Zimeng Luo, Tong Wang et al.ICCV 2023 · 10 citations
- Correlational Image Modeling for Self-Supervised Visual Pre-TrainingWei Li, Jiahao Xie, Chen Change LoyCVPR 2023
- Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation LearningKaiyou Song, Jin Xie, Shan Zhang, Zimeng LuoCVPR 2023
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 183 citations
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong et al.NeurIPS 2021 · 93 citations
- MultiSiam: Self-supervised Multi-instance Siamese Representation Learning for Autonomous DrivingKai Chen, Lanqing Hong, Hang Xu, Zhenguo Li et al.ICCV 2021 · 60 citations
- DetCo: Unsupervised Contrastive Learning for Object DetectionEnze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan et al.ICCV 2021 · 364 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
