Instance Localization for Self-Supervised Detection Pretraining
Ceyuan Yang, Zhirong Wu, Bolei Zhou, Stephen Lin
Abstract
Prior research on self-supervised learning has led to considerable progress on image classification, but often with degraded transfer performance on object detection. The objective of this paper is to advance self-supervised pretrained models specifically for object detection. Based on the inherent difference between classification and detection, we propose a new self-supervised pretext task, called instance localization. Image instances are pasted at various locations and scales onto background images. The pretext task is to predict the instance category given the composited images as well as the foreground bounding boxes. We show that integration of bounding boxes into pretraining promotes better task alignment and architecture alignment for transfer learning. In addition, we propose an augmentation method on the bounding boxes to further enhance the feature alignment. As a result, our model becomes weaker at Imagenet semantic classification but stronger at image patch localization, with an overall stronger pretrained model for object detection. Experimental results demonstrate that our approach yields state-of-the-art transfer learning results for object detection on PASCAL VOC and MSCOCO 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34f6ce55-b47f-4130-a01a-9b4be19e09a2Cited by top-tier papers43
- DetCo: Unsupervised Contrastive Learning for Object DetectionEnze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan et al.ICCV 2021 · 364 citations
- Efficient Self-supervised Vision Transformers for Representation LearningChunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao et al.ICLR 2022 · 228 citations
- MST: Masked Self-Supervised Transformer for Visual RepresentationZhaowen Li, Zhiyang Chen, Fan Yang, Wei Li et al.NeurIPS 2021 · 194 citations
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- VICRegL: Self-Supervised Learning of Local Visual FeaturesAdrien Bardes, Jean Ponce, Yann LeCunNeurIPS 2022 · 189 citations
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
Related papers
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 183 citations
- PreDet: Large-scale weakly supervised pre-training for detectionVignesh Ramanathan, Rui Wang, Dhruv MahajanICCV 2021 · 14 citations
- Siamese DETRZeren Chen, Gengshi Huang, Wei Li, Jianing Teng et al.CVPR 2023
- Aligning Pretraining for Detection via Object-Level Contrastive LearningFangyun Wei, Yue Gao, Zhirong Wu, Han Hu et al.NeurIPS 2021 · 180 citations
- DETReg: Unsupervised Pretraining with Region Priors for Object DetectionAmir Bar, Xin Wang, Vadim Kantorov, Colorado J. Reed et al.CVPR 2022 · 130 citations
