PreDet: Large-scale weakly supervised pre-training for detection
Vignesh Ramanathan, Rui Wang, Dhruv Mahajan
Abstract
State-of-the-art object detection approaches typically rely on pre-trained classification models to achieve better performance and faster convergence. We hypothesize that classification pre-training strives to achieve translation invariance, and consequently ignores the localization aspect of the problem. We propose a new large-scale pre-training strategy for detection, where noisy class labels are available for all images, but not bounding-boxes. In this setting, we augment standard classification pre-training with a new detection-specific pretext task. Motivated by the noise-contrastive learning based self-supervised approaches, we design a task that forces bounding boxes with high-overlap to have similar representations in different views of an image, compared to non-overlapping boxes. We redesign Faster R-CNN modules to perform this task efficiently. Our experimental results show significant improvements over existing weakly-supervised and self-supervised pre-training approaches in both detection accuracy as well as fine-tuning speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b0f1fc7-9b8b-4d66-9bf1-c6fb4c8711caCited by top-tier papers5
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Sylph: A Hypernetwork Framework for Incremental Few-shot Object DetectionLi Yin, Juan M. Perez-Rua, Kevin J. LiangCVPR 2022 · 51 citations
- Proper Reuse of Image Classification Features Improves Object DetectionCristina Nader Vasconcelos, Vighnesh Birodkar, Vincent DumoulinCVPR 2022 · 25 citations
- Efficient Event Camera Data Pretraining with Adaptive Prompt FusionQuanmin Liang, Qiang Li, Shuai Liu, Xinzi Cao et al.ICCV 2025 · 6 citations
- Region-based Cluster Discrimination for Visual Representation LearningYin Xie, Kaicheng Yang, Xiang An, Kun Wu et al.ICCV 2025 · 1 citation
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- DAP: Detection-Aware Pre-Training With Weak SupervisionYuanyi Zhong, Jianfeng Wang, Lijuan Wang, Jian Peng et al.CVPR 2021
- Aligning Pretraining for Detection via Object-Level Contrastive LearningFangyun Wei, Yue Gao, Zhirong Wu, Han Hu et al.NeurIPS 2021 · 180 citations
- Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningYizhou Wang, Meilin Chen, Shixiang Tang, Feng Zhu et al.NeurIPS 2022 · 2 citations
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
