FP-DETR: Detection Transformer Advanced by Fully Pre-training
Wen Wang, Yang Cao, Jing Zhang, Dacheng Tao
Abstract
Large-scale pre-training has proven to be effective for visual representation learning on downstream tasks, especially for improving robustness and generalization. However, the recently developed detection transformers only employ pre-training on its backbone while leaving the key component, i.e., a 12-layer transformer, being trained from scratch, which prevents the model from above benefits. This separated training paradigm is mainly caused by the discrepancy between the upstream and downstream tasks. To mitigate the issue, we propose FP-DETR, a new method that Fully Pre-Trains an encoder-only transformer and smoothly fine-tunes it for object detection via a task adapter. Inspired by the success of textual prompts in NLP, we treat query positional embeddings as visual prompts to help the model attend to the target area (prompting) and recognize the object. To this end, we propose the task adapter which leverages self-attention to model the contextual relation between object query embedding. Experiments on the challenging COCO dataset demonstrate that our FP-DETR achieves competitive performance. Moreover, it enjoys better robustness to common corruptions and generalization to small-size datasets than state-of-the-art detection transformers. Code will be made publicly available at https://github.com/encounter1997/FP-DETR.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7f3a7fb1-0004-44c6-9c34-cb1444d2663bCited by top-tier papers7
- DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.AAAI 2023 · 123 citations
- RU-Net: Regularized Unrolling Network for Scene Graph GenerationXin Lin, Changxing Ding, Jing Zhang, Yibing Zhan et al.CVPR 2022 · 43 citations
- Recurrent Glimpse-based Decoder for Detection with TransformerZhe Chen, Jing Zhang, Dacheng TaoCVPR 2022 · 37 citations
- Sparse Semi-DETR: Sparse Learnable Queries for Semi-Supervised Object DetectionTahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan AfzalCVPR 2024 · 36 citations
- Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual TasksZhiyang Chen, Yousong Zhu, Zhaowen Li, Fan Yang et al.NeurIPS 2022 · 17 citations
Related papers
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang et al.CVPR 2022 · 14 citations
- Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionFeng Liu, Xiaosong Zhang, Zhiliang Peng, Zonghao Guo et al.ICCV 2023 · 30 citations
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 61 citations
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
