UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers
Zhigang Dai, Bolun Cai, Yugeng Lin, Junying Chen
Abstract
Object detection with transformers (DETR) reaches competitive performance with Faster R-CNN via a transformer encoder-decoder architecture. Inspired by the great success of pre-training transformers in natural language processing, we propose a pretext task named random query patch detection to Unsupervisedly Pre-train DETR (UP-DETR) for object detection. Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the original image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we freeze the CNN backbone and propose a patch feature reconstruction branch which is jointly optimized with patch detection. (2) To perform multi-query localization, we introduce UP-DETR from single-query patch and extend it to multiquery patches with object query shuffle and attention mask. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5df47248-8447-4b5d-ac7a-428671b653f9Cited by top-tier papers125
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
Related papers
- FP-DETR: Detection Transformer Advanced by Fully Pre-trainingWen Wang, Yang Cao, Jing Zhang, Dacheng TaoICLR 2022 · 35 citations
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram et al.AAAI 2024 · 4 citations
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang et al.ICCV 2023 · 231 citations
- PreDet: Large-scale weakly supervised pre-training for detectionVignesh Ramanathan, Rui Wang, Dhruv MahajanICCV 2021 · 14 citations
- Learning Dynamic Query Combinations for Transformer-based Object Detection and SegmentationYiming Cui, Linjie Yang, Haichao YuICML 2023 · 13 citations
